ODVS-Bench - a public domain-pricing benchmark
This page shows the engine losing when it loses - and today it does, on at least one slice below. The WITNESSED table below is empty: no submission has been run against the held-out split yet. Every row in the SELF-REPORTED table, including this engine’s own, is scored on the PUBLIC split, whose answers are already public - a signed log there proves only the bytes a submitter chose to upload, never that the answer wasn’t copied from the key. Treat self-reported numbers accordingly.
The public split
- 621 rows: 218 recent held-out sales, 3 curated corpus holdout, 400 ledger sample.
- Every row’s source is on record as redistribution-permitted in this repo’s provenance registry: com-sales, dnjournal, manual-seed, noncom-sales, open-dataset, parkio-results, registry-reports, repo-curated, sedo-manual, sedo-weekly, venue-annual-reports.
- Split hash:
7bfb86fda1478ff509a90c4fca51311e190b85aeb8d611fd4f19bafc32679cf2. A submission log that does not carry this hash is not scored against the current split. - CC-BY-4.0 (rights basis: data/provenance/sources.registry.json, redistribution=permitted rows only)
Witnessed (hidden split)
Run by the maintainer or a steering-group member against a held-out split whose prices were never published anywhere. Only these rows are ever described as a model comparison.
No witnessed submission exists yet. The hidden split itself has not been assembled - it needs rights-cleared, never-published sale prices this repository does not hold.
Self-reported (public split)
Bring-your-own-key runs against the public split above, labelled self-reported because the split’s answers are public.
| Submitter | Model | Coverage | Recent held-out sales within-2x | Curated corpus holdout within-2x | Ledger sample within-2x |
|---|---|---|---|---|---|
| audit.domains (this repo's own estimator) | champion (rev 4.11, generalization) | 621/621 | 31.65% (Wilson 25.84% to 38.1%), n=218 | 0% (Wilson 0% to 56.15%), n=3 | 35.5% (Wilson 30.97% to 40.31%), n=400 |
audit.domains (this repo's own estimator): Self-reported: computed in this repo by the same code the row is about, on the public split above (3 corpus rows survive quarantine of fabricated ground truth - see FIELD_NOTES.md). ACCURACY.md's published headline is measured on the frozen 792, not this split, so the two are not expected to match exactly even where they overlap.
Submitting a result
- Run
tools/odvs-bench/run.tsagainst the public split with your own model endpoint and key - this project spends nothing running your submission. - Score the resulting log with
tools/odvs-bench/score.ts. - Open a pull request adding the row and its log under
selfReportedindata/odvs-bench/leaderboard.json.