Skip to main content
audit.domains
The domain appraisal of record
Accuracy · Revision 4.10 active

Measured against real sales.

How the model is tested - and how wrong it tends to be.

We hold out a fixed share of our price-verified sales - the same domains every run - price them blind, and publish the error for the active methodology and for any candidate revision before it can become the default. A candidate ships only if it beats the incumbent here. Behind it sits a recorded-sales ledger of 459,781 transactions spanning three decades of the aftermarket.

Candidate rev 4.9 (in validation) · like-for-like held-out back-test · as of July 2026
42.9%
Median error
93.1%
Within 3× of sale price
63.2%
Within 2× of sale price

Measured over a fixed held-out cut of price-verified sales, with both models holding the same information set. These are the candidate's numbers; revision 4.10 remains the active default until the candidate is promoted. The full comparison - including the harder blind cut and the retired revision as it shipped - is below.

Hit rate · incumbent vs candidate
Within 2× of sale price+49 pts for candidate
Rev 4.2 (incumbent)
13.8%
Rev 4.9 (candidate)
63.2%
Within 3× of sale price+72 pts for candidate
Rev 4.2 (incumbent)
21.3%
Rev 4.9 (candidate)
93.1%
Held-Out Back-Test · Price-Verified Sales
MetricRev 4.2 (retired)Rev 4.9 blind
own sale withheld
Rev 4.9 like-for-like
same information set
Median error
Median absolute percentage error vs the actual sale price
97%85.5%42.9%
Mean error (MAPE)
Mean absolute percentage error - sensitive to big misses
173%326.8%42.5%
Within 2× of sale price
Share of estimates between half and double the actual price
13.8%28.7%63.2%
Within 3× of sale price
Share of estimates between a third and triple the actual price
21.3%40.2%93.1%
Within ±50%
Share of estimates inside a ±50% band around the sale
13.2%26.4%63.2%
Log-scale RMSE
Root-mean-square error in orders of magnitude (0.30 ≈ 2×)
2.1281.1810.290
Median estimate ÷ sale
1.00 is unbiased; above 1 overprices, below 1 underprices
0.08×0.48×0.57×
Like-for-like (both models holding the same sale data), rev 4.9 leads on 7 of 7 metrics; highlights mark the better model per row. That clean sweep is a like-for-like result only: the “blind” column prices each holdout domain with its own recorded sale withheld - the harder generalization test - and on that cut rev 4.9 trades some ground (its mean error rises to 326.8% against rev 4.2’s 173%) even as it still cuts the median error and log-scale RMSE. It is shown unhighlighted, with no retired-revision counterpart, for transparency. Even with every holdout sale present in its own corpus, rev 4.2 misses by the margins shown - which is why it was retired in favour of rev 4.9. Back-test generated July 18, 2026; re-running the script on the same data reproduces these numbers exactly.
What this means for your number

Trust the published range over the single point estimate: the interval is where the evidence actually places the value. On the like-for-like cut, 63.2% of estimates land within 2× of the recorded sale and 93.1% within 3×, so most numbers sit close to where comparable domains traded. The blind cut - which withholds each domain’s own sale and raises mean error to 326.8% - is the conservative bound to plan against.

How to read this honestly

rev 4.9 (the active methodology) excludes the subject's own recorded sale (generalization cut). The retired rev 4.2 ran as shipped, with the sales corpus and elite price map baked in, so every holdout sale was in its own data and its row is effectively a recall test. The recall cut of rev 4.9 (same information set) is in the committed artifact.

Domain prices are heavy-tailed: a sale can clear at twice or half its fair level on any given day, so even a perfect model would not score 100% here. We publish the factor-2 and factor-3 hit rates because they are the standard a comparable-sales appraisal can actually be held to, and the median ratio because it shows whether the model leans high or low overall.

Two corpus figures appear on this site and they are different things: the full comparable corpus of curated recorded sales that the engine queries for comparables, and the curated, price-verified domain sales used for the back-test - the subset with independently confirmable prices, which is what an evaluation can honestly be scored against. The records scored here are that price-verified back-test subset. Separately, the recorded-sales ledger holds 459,781 recorded transactions kept for research and future calibration; live comparables continue to come from the curated matching corpus until a screened subset passes the published back-test.

Corpus: curated domain-sales database (public aftermarket sales reports and press records); historical sales expansion (verified records, 1995-2025); known-sales register (inflation-adjusted public reports). Evaluation set and split rule are committed to the repository; the full artifact lives at reports/calibration/rev49-backtest.json. Method details: the published methodology →

Put it to the test

The same methodology measured above prices any domain you name - a fair-value range, comparable sales, and a published confidence reading. Your first appraisals are free, no account needed.