New to these reports? Start here
- Dotted-underlined words have a plain-English definition — hover or tap them. Every term is also on the glossary page.
- "Null" means we found nothing, not that something broke. Most reports here are negative results, on purpose — knowing an idea doesn't work is the point.
- Two questions get asked separately. First, is the effect real? Second, is it already priced into the betting odds? An effect can be completely real and still useless to bet on.
- A "calibration" row is a self-check. It runs the same method on something already known to be true. If that fails, the whole report is unreliable — so it's reported alongside the findings.
- If a confidence interval includes zero, the real effect might be nothing at all, so no claim gets made.
Train / dev / holdout: what the accuracy numbers are actually worth
Date: 2026-08-30 · Report: scripts/report_split_accuracy.py (re-runnable)
The discipline
| split | seasons | role |
|---|---|---|
| train | 2021–2023 | feeds the model |
| dev | 2024 | where knobs get chosen — expect it to flatter |
| holdoutholdoutData deliberately set aside and never looked at while developing an idea, then used once at the end as a fair test. Peeking at it first would defeat the purpose. | 2025 | looked at rarely, never tuned against |
| forward | 2026+ | live; the only genuinely untouched evidence |
Defined once in UnifiedAccuracyTracker.SPLIT rather than per-script, because
the failure mode is a script quietly averaging dev into a headline.
Result: the engine is not overfit
All three splits produced by one engine config (81efc25d5d35, verified by the
config_hash now stamped on every graded row):
| split | n | MAEmean absolute errorAverage size of the miss, ignoring direction. If a projection is off by 3 one week and -5 the next, the MAE is 4. Lower is better. | biasbiasWhether the misses lean consistently one way. A projection can have a good average error size but still be biased if it is almost always too high. Bias is often the more fixable problem. | r |
|---|---|---|---|---|
| train | 9,199 | 5.14 | −0.02 | 0.558 |
| dev | 3,114 | 5.19 | −0.36 | 0.573 |
| holdout | 3,110 | 5.18 | +0.01 | 0.553 |
Dev is marginally worse than holdout, and train is barely better. If the knobs had been fitted to noise, dev would flatter and holdout would fall away. It does not happen — the gap is 0.01 points of MAE. That is real evidence the engine generalises, and it is worth more than any single headline figure.
But the 2025 holdout is already spent
Honesty requires naming what informed what:
| decision | contaminates | measured on |
|---|---|---|
| snap/depth re-ranker | dev + holdout | 2025 only |
| volume features | dev + holdout | 2023–2025 pooled |
| player matching (soup kept) | dev + holdout | 2024 + 2025 |
| pace adjustment (rejected) | dev + holdout | 2024 + 2025 |
| weekly rosters | dev + holdout | 2024 + 2025 |
| floor/ceiling calibrationcalibrationA deliberate sanity check on the method itself: run it on something already known to be true. If it fails to detect the known thing, the method is broken and its other results mean nothing. | clean | walk-forwardwalk-forwardEvaluating week by week using only what was knowable before each week, mimicking how the model would actually have been used at the time., prior seasons only |
| TD probability calibration | clean | walk-forward, prior seasons only |
Most knobs were chosen by looking at both 2024 and 2025, so 2025 is no longer a clean holdout for them. The two calibration layers are exempt: they refit on strictly prior seasons at every step, so they never saw the season they scored.
The first genuinely untouched evaluation is 2026 forward, archived at
generation time with provenance='realtime'. Until enough of those accumulate,
5.18 is the best available number, not a clean one — and this document exists so
that distinction survives being quoted.
Guards
config_hashon every graded row;scripts/unit/test_eval_split.pyfails if any split mixes engine versions, or if dev and holdout disagree on config./accuracylabels the season it is showing ("holdout season — already spent: most knobs were measured on it too"), so the page states its own standing.calculate_fantasy_accuracyfilters to one provenance and the dashboard always passes a concrete season, so a pooled figure cannot appear by default.