Skip to main content
audit.domains
A domain appraisal you can check
Accuracy · Revision 4.14 active

Measured against real sales.

How the model is tested - and how wrong it tends to be.

We hold out a fixed set of recorded sales - the same domains every run, never part of the engine’s evidence - price them blind, and publish the error for the active methodology and for any candidate revision before it can become the default. A candidate ships only if it beats the incumbent here. Behind it sits a recorded-sales ledger spanning three decades of the aftermarket.

Two sales figures appear on this site and they are different things: the full comparable corpus of curated recorded sales that the engine queries for comparables, and the held-out back-test set - recorded sales kept out of that corpus, so the engine never sees the prices it is scored against. The headline figures on this page are scored against the 218 recorded sales held out of the engine's evidence for back-testing.

Active: revision 4.14. Structural pricing is revision 4.13; 4.14 adds only the personal-name path - every non-name class is identical to 4.13.

Rev 4.13 (active methodology) · held-out back-test, each sale withheld from the engine's evidence · as of September 2026
85.1%
Median error
39.9%
Within 3× of sale price
29.4%
Within 2× of sale price

Measured over 218 recorded sales held out of the engine's evidence for back-testing, so the engine never saw the prices it is scored against. This is the conservative figure, and the one to plan against. Revision 4.13 is the active structural methodology; revision 4.14 is identical to it for every class except recognized personal names, where it adds a separate path on top. The full comparison against the retired revision, which never saw these sales either, is below.

The same model on three different cuts

The headline is the recent sales cut - sales reported by DNJournal and Sedo, kept out of the corpus the engine draws its evidence from - one of two we measure on the same frozen test set. It was chosen for where its prices come from, not for how the model scores on it, although it is also the cut the model scores best on. The model is weakest on the bulk ledger sample cut. The bulk ledger sample stays published because it is the one closest to screening a drop feed, which is what the numbers get used for. Read the headline as the accuracy on names like the ones in its cut, not as one figure that covers every name.

SliceSales scoredWithin 2× of sale priceMedian errorMedian error as a factor
Recent salesPublished
sales reported by DNJournal and Sedo, kept out of the corpus the engine draws its evidence from
21829.4%
23.7% - 35.7%
85.1%
80.6% - 89.0%
5.0×
3.5× - 6.5×
Bulk ledger sample
a deterministic sample of the recorded-sales ledger: the auction-heavy bulk distribution a drop list is drawn from
40013.0%
10.1% - 16.7%
871.0%
720.6% - 1,108.9%
11.3×
9.5× - 13.2×

One estimator (champion (rev 4.13, generalization)), one frozen leakage-guarded test set of 618 names, two disjoint slices; no name is scored twice. The published row is the same measurement as the headline above. Test set 5866aaffbc4e, scored in tools/scoreboard/baselines/champion.json. The curated holdout cut is no longer scored: the 2026-09-23 provenance audit withdrew 171 of its sales as generated or unrecorded, leaving 3, too few to measure.

The smaller figure under each number is its 95% interval - the range the sample size can actually support. Read the interval, not the decimal place.

Median error and median error as a factor measure the same misses two ways, and the second is the one to trust. A percentage error cannot go below -100% but has no ceiling above, so on prices - which move by multiples, not by percentages - the same miss scores twice as badly when the estimate lands high as when it lands low. A 2× underestimate scores 50%; a 2× overestimate scores 100%; a 10× underestimate scores 90%, which reads better than the mild 2× overestimate it is far worse than. The factor column removes that tilt by measuring the miss in whichever direction it falls, so the published 5.0× means half of all estimates land within 5.0× of the eventual sale price - above or below.

How well it orders

Every figure above asks how close a price is. This asks a different question: given a list, does the engine put the right names at the top? For anyone triaging a drop list or a portfolio, that ordering is the thing being bought, and it is measured on the same frozen test set, slice by slice.

Ordering quality of the shipped rank model, by slice
SliceNamesRank correlationTop 50 correctValue in top 50
Recent sales2180.35138%53.11%
Bulk ledger sample4000.33542%36.72%

Rank correlation is Spearman's: 1.00 would be a perfect ordering, 0.00 no better than shuffling. Top 50 correct is how many of the fifty names the model puts first really do belong in the slice's top fifty. Value in top 50 is the share of everything that slice eventually sold for which those fifty names account for - the figure that decides whether triaging by this order is worth doing.

The slices are reported separately and not pooled. A pooled figure would read higher than any row here, because the slices differ in price by orders of magnitude and most of a pooled ordering is decided by which slice a name came from rather than by any judgement about the name. Scored in tools/scoreboard/baselines/two-output.json against test set 5866aaffbc4e, the same one the price table above uses.

Against the retired revision, on sales neither revision had seen - each sale withheld from the engine's evidence. The same figures as the headline above.
Hit rate · incumbent vs candidate
Within 2× of sale price+11 pts for candidate
Rev 4.2 (incumbent)
18.8%
Rev 4.13 (candidate)
29.4%
Within 3× of sale price+14 pts for candidate
Rev 4.2 (incumbent)
26.1%
Rev 4.13 (candidate)
39.9%
Held-Out Back-Test · Recorded Sales
MetricRev 4.2 (retired)Rev 4.13 (active)
sale never seen
Rev 4.13 recall
the same sales with each one's own sale added to the engine's evidence
Median error
Median absolute percentage error vs the actual sale price
99.8%85.1%26%
Mean error (MAPE)
Mean absolute percentage error - sensitive to big misses
1513.5%96.5%28%
Within 2× of sale price
Share of estimates between half and double the actual price
18.8%29.4%94%
Within 3× of sale price
Share of estimates between a third and triple the actual price
26.1%39.9%99.1%
Within ±50%
Share of estimates inside a ±50% band around the sale
15.1%23.9%92.7%
Log-scale RMSE
Root-mean-square error in orders of magnitude (0.30 ≈ 2×)
1.6640.9760.171
Median estimate ÷ sale
1.00 is unbiased; above 1 overprices, below 1 underprices
0.72×0.25×0.76×
Neither revision’s evidence holds these sales, so the two revision columns are the same test: rev 4.13 leads on 6 of 7 metrics, and highlights mark the better model per row. The last column adds each sale to rev 4.13’s own evidence, to show what the engine does when it already holds the answer; it has no retired-revision counterpart and is shown unhighlighted. Back-test generated September 30, 2026; re-running the script on the same data reproduces these numbers exactly.
What this means for your number

Trust the published range over the single point estimate: the interval is where the evidence actually places the value. With each sale withheld from the engine’s evidence, 29.4% of estimates land within 2× of the recorded sale and 39.9% within 3×, so the range carries the answer and the point estimate is only its midpoint. If the engine already held the sale - the same sales with each one's own sale added to the engine's evidence - those read 94% and 99.1%, which is exactly why the withheld-sale figures are the ones published above.

How to read this honestly

None of these sales is in the corpus the engine draws its evidence from. Rev 4.13 (the active methodology) is scored blind: each subject is also excluded from its comparable pool. The retired rev 4.2 ran as shipped and never saw these sales either - none is in its sales corpus or elite price map - so its row is a BLIND comparison and pairs with rev 4.13's blind cut, not the recall cut. The recall cut adds each subject's own sale to rev 4.13's evidence, to show what the engine does when it already holds the answer.

Domain prices are heavy-tailed: a sale can clear at twice or half its fair level on any given day, so even a perfect model would not score 100% here. We publish the factor-2 and factor-3 hit rates because they are the standard a comparable-sales appraisal can actually be held to, and the median ratio because it shows whether the model leans high or low overall.

Two sales figures appear on this site and they are different things: the full comparable corpus of curated recorded sales that the engine queries for comparables, and the held-out back-test set - recorded sales kept out of that corpus, so the engine never sees the prices it is scored against. The records scored here are that held-out back-test set.

Corpus: recorded sales dated 1999-2026; DNJournal weekly sales reports; curated famous sales and trade-press reports, each cited to the public report that records it. Evaluation set and split rule are committed to the repository. Revision 4.14 adds only the personal-name path, which this back-test does not exercise, so this revision 4.13 back-test is the accuracy of 4.14 for every class except recognized personal names. Method details: the published methodology →

Revision history

How accurate is the audit.domains appraisal, and did the methodology improve?

Measured on 218 recorded sales, each sale withheld from the engine's evidence, evaluated September 30, 2026: the live methodology (revision 4.14) has a median error of 85.1% and lands within 2× of the recorded price 29.4% of the time, within 3× 39.9% of the time. The revision it replaced (4.2, run as shipped, never having seen these sales, the same 218 sales) had a median error of 99.8% and was within 2× 18.8% of the time. Every revision is below, including the two that carry no figure because they have not been measured.

RevisionStatusSales scoredMedian errorWithin 2×Within 3×
4.2
run as shipped, never having seen these sales
Retired as default21899.8%18.8%26.1%
4.13
each sale withheld from the engine's evidence
Active structural base21885.1%29.4%39.9%
4.14
adds a personal-name path on top of revision 4.13; identical figures for every class except recognized personal names
Active default21885.1%29.4%39.9%
5.0
no published back-test
Flag off----
5.1
no published back-test
Flag off----

Machine-readable copy: /accuracy/backtest.json carries this table with each figure’s cut label and basis attached, generated from the same committed facts as this page, and fetchable from any origin.

Put it to the test

The same methodology measured above prices any domain you name - a fair-value range, comparable sales, and the evidence behind the figure. Your first appraisals are free, no account needed.