How we appraise a domain.
Every factor the model counts, and what each is worth.
No black boxes. Every appraisal is the product of eight weighted factors, disciplined against 933 curated real recorded sales, and bounded by an explicit confidence interval. The same domain always returns the same figure under the same methodology revision - this page sets out exactly what we measure and how the pieces fit together.
Accuracy is measured, not asserted: see the held-out back-test metrics →
Every report we issue is reconstructible from its inputs: the same domain always returns the same figure under the same methodology revision, and each appraisal records the revision that produced it. Issued certificates keep the revision they were issued under - a model update never restates an existing file.
Revision 4.10 is the revision pricing new appraisals right now, having replaced revision 4.2 as the default. Revision 4.9 scores the eight factors below, prices them through a log-space structural model, then blends in an anchor built from matched real sales - each sale first indexed to the appraised name (adjusted for zone, length, keyword and brandability differences), then combined as a similarity- and recency-weighted median, so a comparable contributes the market premium it demonstrated rather than its raw price level. Revision 4.2 scores length on a linear curve and reads keyword demand from aggregate engine buckets; 4.9 replaces those with a convex scarcity curve and the cost-per-click dataset described below. Revision 4.10 adds a personal-name pricing path so a recognized first name is valued as the liquid end-user .com class it is - on a popularity-and-length curve floored by tier - instead of being garbage-floored or priced off unrelated comparables; every other class is identical to revision 4.9. A candidate becomes the default only after beating the incumbent on the held-out back-test.
Length is scored on a convex scarcity curve: one-to-three-character names are an order of magnitude scarcer and score far higher than mid-length names, with the curve steepest at the short end and flattening toward a floor past roughly a dozen characters. Shortness is rewarded the way real sales reward it - not linearly.
Comparables are drawn exclusively from the comparable corpus of curated recorded sales - never from generated or modeled “sales”. Each candidate gets one continuous similarity score that weights shared tokens and word family, zone relationship, spelling overlap, structural class, industry overlap, and length proximity, with shared tokens and word-family kinship carrying the most influence; spelling overlap between two unrelated dictionary words is deliberately discounted, because words trade on meaning, not spelling.
Every listed sale is then indexed to the appraised name: its price is adjusted, on the same published structural model, for the differences between it and the subject - zone, length, keyword demand, brandability - and the indexed figure is what the anchor consumes. A same-name sale in another zone translates through the documented zone offsets instead of exporting its raw dollar figure, and recent sales carry more anchor weight than decades-old prints under a fixed corpus-edition clock.
Out-of-class premiums are rejected by a robust median-deviation screen on the indexed scale unless the match is near-exact, and weak matches below a minimum similarity are never listed. When too few strong matches exist, the report says so and the confidence reading drops - the table is never padded. A recorded prior sale of the appraised domain itself is labelled and anchors the estimate directly.
Data sources: curated domain-sales database (public aftermarket sales reports and press records); historical sales expansion (verified records, 1995-2025); known-sales register (inflation-adjusted public reports). Keyword economics come from the bundled cost-per-click dataset (46 verticals, id bundled-cpc-2024-v1); a provider interface allows a licensed live feed to replace it without touching scoring code. Sale records outside $100 - $50,000,000 are screened out as non-domain transactions.
- The personal-name model itself (detection, popularity curve, cohort discipline, tier floors) is identical to revision 4.8.
- Renumbered because the class gating underneath changed in revision 4.9: the revision stamp scopes the valuation cache, so a result priced under the corrected gating must never be served under a pre-correction stamp.
- Pronounceable brandable names (measured phonotactic gate: vowel balance, no long consonant runs, an early vowel) are exempt from the garbage-cap ladder. The garbage scorer's random-string signals mis-read this class - measured against the real-sales corpus, 28.4% of genuine recorded transactions would have been capped to the junk floor; the gate salvages the class while keyboard-mash names and confident brand typo-squats remain capped by their own signals.
- Eligibility for the comparables overlay is decided by name class instead of by the incoming value: a legitimate name a protection layer mis-capped to a junk price is repriced by the overlay, while genuine garbage-class strings never reach the comparable anchor at any value.
- Hyphenated all-numeric names take the pure-digit class ceiling on their stripped digits (123-456 prices as the discounted form of 123456), closing the gap where they escaped every numeric ceiling by classifying as hyphenated.
- The personal-name model itself (detection, popularity curve, cohort discipline, tier floors) is identical to revision 4.6.
- Renumbered because the base comparable machinery underneath changed in revision 4.7: the revision stamp scopes the valuation cache, so an indexed-anchor result must never be served under a pre-indexing stamp.
- Each comparable's price is indexed to the subject on the published structural model - adjusted for zone, length, keyword demand and brandability differences - and the anchor is the similarity- and recency-weighted median of those indexed prices. A same-name sale in another zone now translates through the documented zone offsets instead of exporting its raw dollar figure.
- Premiums transfer 1:1 up to two-and-a-half orders of magnitude beyond typical (the p95 of recorded premiums); only headline-grade deviations damp, because the held-out back-test rejected every whole-range compression.
- Recent sales carry more anchor weight than decades-old prints: a ten-year half-life with a floor, on the fixed corpus-edition clock - never the wall clock, so determinism is unchanged.
- The outlier screen and the dispersion that drives confidence both moved to the indexed scale: raw price spread across zones is expected and explained; residual disagreement is genuine uncertainty.
- Similarity adds a word-family channel (testing / tester / tests are one market), damps spelling coincidence between distinct dictionary words and across structural classes, and scales dictionary-class matches by word-tier proximity.
- Elite and known-entity reports, whose price is owned by the protection layers, now carry display-only methodology-scored comparables instead of an empty evidence panel; each report row and the analyst note publish the per-sale indexed figure.
- The personal-name model itself (detection, popularity curve, cohort discipline, tier floors) is identical to revision 4.4.
- Renumbered because the base structural path underneath changed in revision 4.5: the revision stamp scopes the valuation cache, so a recalibrated result must never be served under a pre-recalibration stamp.
- The structural intercept rises by 0.3 log10 (2x): the rev 4.3 structural prior predicted a median 30% of the recorded price on held-out sales; slope coefficients are unchanged so relative ranking between names is preserved.
- The comparable anchor may now carry up to 95% of the final log-price (was 84%): on comp-rich names the residual structural share was dragging estimates roughly 2x under their own matched-sales anchor.
- The correction is deliberately partial: a larger lift, a relaxed thin-evidence gate, and a full least-squares refit each scored better still on the holdout but were rejected - they inflate speculative thin-evidence compounds out of their golden-set band, or collapse the length slope toward the corpus median (a survivorship artifact of a notable-sales corpus).
- Recognized first names (a bundled, versioned given-name dataset with a popularity weight) are detected as a distinct value class - they are not dictionary words or keyword terms, so the structural model and comp matcher could not read them.
- First names are priced on a small log-space curve in popularity and length, disciplined by a personal-name comparable cohort when one exists and floored by popularity tier, so a recognized name can never collapse to a near-zero value (aubree.com no longer reads ~$15).
- The comparable matcher restricts a first name's cohort to other recognized first names, so an incidental, unrelated sale is never shown as its comparable; when no first-name sale exists in the curated corpus the report discloses the thin evidence with a wide band.
- Every domain outside the personal-name class is priced exactly as in revision 4.3 and keeps the 4.3 stamp; the curated back-test corpus contains no first-name sales, so held-out accuracy is unchanged from 4.3.
- Length follows a convex scarcity curve - one-to-three character names price an order of magnitude above mid-length names - replacing the rev 4.2 linear curve.
- Every TLD zone carries its own market multiplier and liquidity weight, fit from recorded sales, instead of a single shared discount.
- Keyword demand and search footprint read advertiser cost-per-click and monthly search volume from the bundled CPC dataset, not aggregate engine buckets.
- The structural estimate is blended in log space with a similarity-weighted anchor from matched real sales; a recorded prior sale of the domain itself anchors the estimate directly.
- Back-tested with a generalization cut that excludes the subject's own recorded sale, so the held-out numbers measure prediction rather than recall.
- Length scored linearly in character count - 100 * (12 - n) / 12, floored.
- Keyword demand and search footprint came from the engine's aggregate market and technical buckets.
- No similarity-weighted comparable-sales anchor; estimates leaned on the structural model alone.