New to these reports? Start here
  • Dotted-underlined words have a plain-English definition — hover or tap them. Every term is also on the glossary page.
  • "Null" means we found nothing, not that something broke. Most reports here are negative results, on purpose — knowing an idea doesn't work is the point.
  • Two questions get asked separately. First, is the effect real? Second, is it already priced into the betting odds? An effect can be completely real and still useless to bet on.
  • A "calibration" row is a self-check. It runs the same method on something already known to be true. If that fails, the whole report is unreliable — so it's reported alongside the findings.
  • If a confidence interval includes zero, the real effect might be nothing at all, so no claim gets made.

Train / dev / holdout: what the accuracy numbers are actually worth

Date: 2026-08-30 · Report: scripts/report_split_accuracy.py (re-runnable)

The discipline

split seasons role
train 2021–2023 feeds the model
dev 2024 where knobs get chosen — expect it to flatter
holdoutholdoutData deliberately set aside and never looked at while developing an idea, then used once at the end as a fair test. Peeking at it first would defeat the purpose. 2025 looked at rarely, never tuned against
forward 2026+ live; the only genuinely untouched evidence

Defined once in UnifiedAccuracyTracker.SPLIT rather than per-script, because the failure mode is a script quietly averaging dev into a headline.

Result: the engine is not overfit

All three splits produced by one engine config (81efc25d5d35, verified by the config_hash now stamped on every graded row):

split n MAEmean absolute errorAverage size of the miss, ignoring direction. If a projection is off by 3 one week and -5 the next, the MAE is 4. Lower is better. biasbiasWhether the misses lean consistently one way. A projection can have a good average error size but still be biased if it is almost always too high. Bias is often the more fixable problem. r
train 9,199 5.14 −0.02 0.558
dev 3,114 5.19 −0.36 0.573
holdout 3,110 5.18 +0.01 0.553

Dev is marginally worse than holdout, and train is barely better. If the knobs had been fitted to noise, dev would flatter and holdout would fall away. It does not happen — the gap is 0.01 points of MAE. That is real evidence the engine generalises, and it is worth more than any single headline figure.

But the 2025 holdout is already spent

Honesty requires naming what informed what:

decision contaminates measured on
snap/depth re-ranker dev + holdout 2025 only
volume features dev + holdout 2023–2025 pooled
player matching (soup kept) dev + holdout 2024 + 2025
pace adjustment (rejected) dev + holdout 2024 + 2025
weekly rosters dev + holdout 2024 + 2025
floor/ceiling calibrationcalibrationA deliberate sanity check on the method itself: run it on something already known to be true. If it fails to detect the known thing, the method is broken and its other results mean nothing. clean walk-forwardwalk-forwardEvaluating week by week using only what was knowable before each week, mimicking how the model would actually have been used at the time., prior seasons only
TD probability calibration clean walk-forward, prior seasons only

Most knobs were chosen by looking at both 2024 and 2025, so 2025 is no longer a clean holdout for them. The two calibration layers are exempt: they refit on strictly prior seasons at every step, so they never saw the season they scored.

The first genuinely untouched evaluation is 2026 forward, archived at generation time with provenance='realtime'. Until enough of those accumulate, 5.18 is the best available number, not a clean one — and this document exists so that distinction survives being quoted.

Guards

  • config_hash on every graded row; scripts/unit/test_eval_split.py fails if any split mixes engine versions, or if dev and holdout disagree on config.
  • /accuracy labels the season it is showing ("holdout season — already spent: most knobs were measured on it too"), so the page states its own standing.
  • calculate_fantasy_accuracy filters to one provenance and the dashboard always passes a concrete season, so a pooled figure cannot appear by default.