New to these reports? Start here
- Dotted-underlined words have a plain-English definition — hover or tap them. Every term is also on the glossary page.
- "Null" means we found nothing, not that something broke. Most reports here are negative results, on purpose — knowing an idea doesn't work is the point.
- Two questions get asked separately. First, is the effect real? Second, is it already priced into the betting odds? An effect can be completely real and still useless to bet on.
- A "calibration" row is a self-check. It runs the same method on something already known to be true. If that fails, the whole report is unreliable — so it's reported alongside the findings.
- If a confidence interval includes zero, the real effect might be nothing at all, so no claim gets made.
Backup QB: shade the receivers, not the quarterback
Date: 2026-08-31 · Verdict: ✅ SHIPPED, receiving only
The measurement
backtests/nfl_context_signals established that a team starting a backup loses
−3.41 points [−4.88, −1.96]. That is a team-scoring number, so it had to be
converted into per-stat multipliers before the engine could use it. Measured
directly on 2019–2025, using the same starter definition (≥85% of team attempts
in both prior weeks) — 775 backup team-weeks against 2,217 starter weeks:
| stat | ratio | 95% CI95% confidence intervalThe range the true value is plausibly in. If this range includes zero, we cannot rule out that the real effect is nothing at all. |
|---|---|---|
| passing_tds | 0.759 | [0.710, 0.814] |
| passing_yards | 0.896 | [0.874, 0.920] |
| receiving_yards | 0.896 | [0.873, 0.920] |
| receptions | 0.932 | [0.913, 0.953] |
| targets | 0.960 | [0.941, 0.980] |
| attempts | 0.962 | [0.942, 0.982] |
| interceptions | 1.223 | [1.113, 1.344] |
| carries | 0.995 | [0.972, 1.018] — spans 1 |
| rushing_yards | 0.969 | [0.935, 1.004] — spans 1 |
Teams do not run more behind a backup. They pass worse. Both rushing ratios span 1.0 and are left alone.
Applying it to the QB was wrong, and the A/B said so
The obvious implementation — apply every ratio, including the passing ones — made things worse in both seasons:
| 2024 dev | 2025 holdoutholdoutData deliberately set aside and never looked at while developing an idea, then used once at the end as a fair test. Peeking at it first would defeat the purpose. | |
|---|---|---|
| passing_yards MAEmean absolute errorAverage size of the miss, ignoring direction. If a projection is off by 3 one week and -5 the next, the MAE is 4. Lower is better. | +1.068 | +0.650 |
| overall | 6 worse / 3 better | 3 worse / 6 better |
The reason is a double-count: a backup's own recency-weighted baseline is built from his games, so it already encodes that he is a backup. Shading it again applies the discount twice. His pass-catchers are the opposite case — their baselines were built with the starter throwing, and nothing else in the engine knows the thrower changed.
Receiving-only ships
| 2024 dev | 2025 holdout | |
|---|---|---|
| receiving_yards | 19.510 → 19.504 | 18.540 → 18.486 |
| receiving_tds | 0.289 → 0.285 | 0.287 → 0.282 |
| receptions | unchanged | 1.421 → 1.416 |
| targets | +0.001 | 1.825 → 1.822 |
| overall | target categories improve | all six moved categories improve |
Small, but it improves the categories it targets in both seasons, and the mechanism was measured rather than assumed.
⚠️ Noted while reading the results: changing a receiving stat moves
rushing slightly too (±0.01), because _top_end_shrink_ratio keys on total
PPRpoints per receptionA fantasy scoring format that awards a point for every catch, which raises the value of high-volume receivers. and rescales the whole stat line. Any per-stat A/B on this engine will show
small movement in categories it did not touch — that is coupling, not noise, and
it is worth remembering before reading a ±0.01 shift as signal.
Scope
Detection uses prior weeks only (the two weeks before the one being projected) and returns "not a backup" whenever both are unavailable, so a season opener or a data gap can never manufacture a discount. ~24% of team-weeks qualify, which matches the 26% in the historical sample — the definition is deliberately broad and also catches a starter's first game back.