Measured against real sales.
How the model is tested - and how wrong it tends to be.
We hold out a fixed share of our price-verified sales - the same domains every run - price them blind, and publish the error for the active methodology and for any candidate revision before it can become the default. A candidate ships only if it beats the incumbent here. Behind it sits a recorded-sales ledger of 459,781 transactions spanning three decades of the aftermarket.
Measured over a fixed held-out cut of price-verified sales, with both models holding the same information set. These are the candidate's numbers; revision 4.10 remains the active default until the candidate is promoted. The full comparison - including the harder blind cut and the retired revision as it shipped - is below.
| Metric | Rev 4.2 (retired) | Rev 4.9 blind own sale withheld | Rev 4.9 like-for-like same information set |
|---|---|---|---|
Median error Median absolute percentage error vs the actual sale price | 97% | 85.5% | 42.9% |
Mean error (MAPE) Mean absolute percentage error - sensitive to big misses | 173% | 326.8% | 42.5% |
Within 2× of sale price Share of estimates between half and double the actual price | 13.8% | 28.7% | 63.2% |
Within 3× of sale price Share of estimates between a third and triple the actual price | 21.3% | 40.2% | 93.1% |
Within ±50% Share of estimates inside a ±50% band around the sale | 13.2% | 26.4% | 63.2% |
Log-scale RMSE Root-mean-square error in orders of magnitude (0.30 ≈ 2×) | 2.128 | 1.181 | 0.290 |
Median estimate ÷ sale 1.00 is unbiased; above 1 overprices, below 1 underprices | 0.08× | 0.48× | 0.57× |
Trust the published range over the single point estimate: the interval is where the evidence actually places the value. On the like-for-like cut, 63.2% of estimates land within 2× of the recorded sale and 93.1% within 3×, so most numbers sit close to where comparable domains traded. The blind cut - which withholds each domain’s own sale and raises mean error to 326.8% - is the conservative bound to plan against.
rev 4.9 (the active methodology) excludes the subject's own recorded sale (generalization cut). The retired rev 4.2 ran as shipped, with the sales corpus and elite price map baked in, so every holdout sale was in its own data and its row is effectively a recall test. The recall cut of rev 4.9 (same information set) is in the committed artifact.
Domain prices are heavy-tailed: a sale can clear at twice or half its fair level on any given day, so even a perfect model would not score 100% here. We publish the factor-2 and factor-3 hit rates because they are the standard a comparable-sales appraisal can actually be held to, and the median ratio because it shows whether the model leans high or low overall.
Two corpus figures appear on this site and they are different things: the full comparable corpus of curated recorded sales that the engine queries for comparables, and the curated, price-verified domain sales used for the back-test - the subset with independently confirmable prices, which is what an evaluation can honestly be scored against. The records scored here are that price-verified back-test subset. Separately, the recorded-sales ledger holds 459,781 recorded transactions kept for research and future calibration; live comparables continue to come from the curated matching corpus until a screened subset passes the published back-test.
Corpus: curated domain-sales database (public aftermarket sales reports and press records); historical sales expansion (verified records, 1995-2025); known-sales register (inflation-adjusted public reports). Evaluation set and split rule are committed to the repository; the full artifact lives at reports/calibration/rev49-backtest.json. Method details: the published methodology →
The same methodology measured above prices any domain you name - a fair-value range, comparable sales, and a published confidence reading. Your first appraisals are free, no account needed.