New to these reports? Start here
- Dotted-underlined words have a plain-English definition — hover or tap them. Every term is also on the glossary page.
- "Null" means we found nothing, not that something broke. Most reports here are negative results, on purpose — knowing an idea doesn't work is the point.
- Two questions get asked separately. First, is the effect real? Second, is it already priced into the betting odds? An effect can be completely real and still useless to bet on.
- A "calibration" row is a self-check. It runs the same method on something already known to be true. If that fails, the whole report is unreliable — so it's reported alongside the findings.
- If a confidence interval includes zero, the real effect might be nothing at all, so no claim gets made.
Fantasy/prop model audit — what is measurable, what is predictable, what to build
2026-08-30. Four parallel investigations: published literature, evaluation methodology, current engine features, and available data. Verdict up front:
The models are not being measured against the right target, and roughly a third of the stats being modelled are known-unpredictable. Both facts were invisible because the evaluation had no baseline that could reveal them.
1. The target is corrupt — fix before anything else
unified_accuracy_tracker.py grades against stats.get('interceptions').
nflreadpy's column is passing_interceptions; no bare interceptions
column exists, so the lookup silently returns nothing.
SUM(actual_interceptions) over all 15,979 graded rows = 0.0
SUM(actual_passing_tds) over the same rows = 2,774
Every stored "actual PPRpoints per receptionA fantasy scoring format that awards a point for every catch, which raises the value of high-volume receivers." is therefore not PPR — missing the interception
penalty (~1.4 pts/QB-game) and, from the same bug class, two of three fumble
categories (only sack_fumbles_lost is subtracted). The projection does
include an INT penalty, so measured QB biasbiasWhether the misses lean consistently one way. A projection can have a good average error size but still be biased if it is almost always too high. Bias is often the more fixable problem. reads −0.60 (under-projecting) when
the truth is roughly +0.8 (over-projecting) — a sign flip.
Every accuracy number this project has ever published is measured against this target. Re-grade all 15,979 rows before drawing any conclusion from them.
2. TD models are worse than predicting zero — and that is CORRECT
| stat | n | model MAEmean absolute errorAverage size of the miss, ignoring direction. If a projection is off by 3 one week and -5 the next, the MAE is 4. Lower is better. | always-0 |
|---|---|---|---|
| rushing_tds | 6,010 | 0.366 | 0.268 |
| receiving_tds | 14,073 | 0.288 | 0.190 |
36% and 52% worse than a constant. The literature says this is not a bug to fix but a signal not to chase: TD rate year-over-year r = 0.19 (WR), 0.05 (RB rushing). Touchdown rate is close to noise; only the volume that produces it is stable.
MAE_REPORT.md already used the "zero is MAE-optimal" argument to reject
fumbles-from-data, then never applied it to TDs — because the only comparator in
the harness was a trailing mean that is also worse than zero.
3. What the literature says is predictable
Volume is stable; efficiency is noise. The most replicated result in the field, and it maps directly onto what to build.
| signal | y/y r | verdict |
|---|---|---|
| Targets/game (WR) | 0.70 | build on it |
| Target share / WOPR / air-yards share | >0.70 | build on it |
| Receptions, receiving yds per game | 0.67–0.70 | build on it |
| Routes/game (RB) | 0.64 | build on it |
| TPRR / YPRR | 0.39–0.64 | usable |
| EPAexpected points addedHow much a play changed the number of points its team should expect to score on that drive. A better measure of a play than raw yards./attempt | 0.18–0.22 | noise |
| TD rate | 0.05–0.19 | noise |
| Success rate | 0.01 | noise |
| Contested-catch rate | 0.02 | noise |
RB weighted opportunities → fantasy points/game R² = 0.82, while no efficiency metric cleared R² = 0.15 (n=770, 2017–21).
Opponent adjustment: damping toward zero is right. Per one rank of opponent defense, weekly: QB −0.07 PPG, RB −0.13, WR −0.09, TE not significant. Own offense is 2–4× larger. Fantasy-points-allowed y/y r is only 0.16–0.27, and Open Source Football found opponent-adjusting defensive EPA reduces predictive power (0.6404 vs 0.6348 — inside noise). The engine's damped defensive multiplier is already correct; do not revisit it.
Game script has the one real published environment coefficient:
Pass Ratio = 58.76% − 0.71 × Game Script (n=512 team-games). Pass rate 66%
when trailing vs 49% when leading; a target is worth 2.64× a carry in PPR for
RBs. This is distinct from the Vegas total the engine already uses.
Vegas implied team total → player points has no credible published coefficient. The often-cited ">0.80" is percentile-bucketed, QB-only, no sample period. The market is an unbiased predictor of margin — a baseline to match, not an edge to harvest.
4. Evaluation defects beyond the target
Ranked, from the methodology audit:
- Corrupt target (§1).
- TD stats scored only against a baseline that is itself worse than zero (§2).
accuracy_within_5/10is structurally meaningless for counts — every TD stat reads 100.0% because the range is 0–5 and the band is ±5.- MAPE and
within_20pctdivide by actual and drop non-finite rows, deleting exactly the 1,378 zero-actual rows — the hardest cases. MNAR, flattering by construction. Same shape as theclv_unreliablebug in the betting stack. - Survivorship: the player universe comes from
load_rosters(max(season))filtered tostatus == 'ACT', so anyone who finished on IR/PUP or was cut is absent from every week including the ones he played — and season-ending injury correlates with poor production beforehand. - Backfilled rows run without the injury gate; realtime rows include it. Two estimators pooled and reported as the deployed one.
- Dev season (2024) pooled into headline numbers after 12 config variants were tried on it. Several kept deltas (0.02–0.05) are inside the noise of a 12-way search.
- No uncertainty anywhere:
AccuracyMetricscarriessample_sizeand no interval. Every "improvement" in MAE_REPORT is a bare point estimate. - 21.5% of projections are never graded, and the dashboard reports the joined count as sample size — so an availability regression that stopped projecting fringe players would lower MAE and look like progress.
Is model 5.20 vs baseline 5.30 real? Clustered by week, yes: paired difference −0.248, 95% block-bootstrap CI [−0.297, −0.198]. Clustered by player it vanishes: −0.030, t = −0.50. The gain is concentrated in high-volume starters and is absent for everyone else — a much weaker claim than "beats persistence", and both numbers are computed against the corrupt target.
5. Data available and unused
The 2021 floor is artificial. nflreadpy resolves to 1999 for player
stats, schedules, rosters and pbp (verified by call). Already local:
nfl_stats.db::nfl_player_game_stats— 39,385 player-weeks, 2019–2025, already carrying target_share, air_yards_share and EPA: precisely the volume metrics the literature ranks highest, and the engine consumes none of them.schedules_1999_2025.feather— 7,276 games with closing spread/total.nfl_weather.db— 2,895 games, 2015–2025.ff_ecr.db— 324,378 FantasyPros consensus snapshots, 2019–2026. This is the benchmark the evaluation should be measured against; beating a trailing average is table stakes.
Loaders exist but persist nothing for snap counts (2012+), NGS (2016+), depth charts (2001+), injuries (2009+), FTN (2022+).
⚠ Corrupt: stat_props.db receptions lines have a 64.5% push rate because
the line column holds whole integers, not half-point book lines — wrong field,
same shape as nba_sbr_odds.close_home_spread. player_props.player_id is
100% NULLnull resultA test that found nothing. "Null" is the starting assumption that there is no real effect; a "null result" means the data gave us no reason to abandon that assumption. It does not mean the data was missing or the test failed to run.. Eight .db files whose names imply live pipelines are empty.
6. Discrepancy worth checking
The engine sets Doubtful → 0.01 on the basis that such players play 0.7% of the time. Published rates (2017–23, n>2,000) are Doubtful 5.9%, Questionable 71%. The Questionable tier matches well; Doubtful is ~8× off. Likely a sample or label-timing artifact in our derivation, but it should be re-derived.
7. Order of work
- Fix the target (INT column, fumble categories), re-grade 15,979 rows.
- Rebuild evaluation: baselines = constant / persistence / depth-rank / ECRexpert consensus rankingFantasyPros' averaged ranking across many fantasy analysts. A strong benchmark, and hard to beat.; skill scores; paired CIs clustered by week and player; delete within-5, MAPE, within-20pct; report coverage; separate dev from holdoutholdoutData deliberately set aside and never looked at while developing an idea, then used once at the end as a fair test. Peeking at it first would defeat the purpose..
- Stop point-estimating TDs and any efficiency feature. Score TDs as P(≥1) on a distribution instead — the engine already emits floor/ceiling and is never scored on it.
- Build volume: target share / air-yards share / routes from data already on disk, plus game script → pass ratio.
- Then extend history to 1999 and re-measure.
Open questions the literature has never answered
No peer-reviewed weekly NFL projection benchmark exists. Nobody has published CRPS/calibrationcalibrationA deliberate sanity check on the method itself: run it on something already known to be true. If it fails to detect the known thing, the method is broken and its other results mean nothing. for fantasy distributions, a stat-decomposition vs direct-FP head-to-head, weekly stabilisation points for usage metrics, or a correlationcorrelationHow closely two things move together, from -1 (opposite) through 0 (unrelated) to +1 (in lockstep). It does not by itself mean one causes the other. between projection MAE and actual lineup win rate. All four are measurable with the data on this disk, and would be novel rather than replication.