New to these reports? Start here
- Dotted-underlined words have a plain-English definition — hover or tap them. Every term is also on the glossary page.
- "Null" means we found nothing, not that something broke. Most reports here are negative results, on purpose — knowing an idea doesn't work is the point.
- Two questions get asked separately. First, is the effect real? Second, is it already priced into the betting odds? An effect can be completely real and still useless to bet on.
- A "calibration" row is a self-check. It runs the same method on something already known to be true. If that fails, the whole report is unreliable — so it's reported alongside the findings.
- If a confidence interval includes zero, the real effect might be nothing at all, so no claim gets made.
The team+position "soup" is not a bug — it beats the clean matcher
Date: 2026-08-30
Harness: scripts/evaluate_fantasy_mae.py --set unified_player_matching=true|false
Verdict: ❌ REJECTED — keep unified_player_matching = False
What prompted it
shrinkage_positions is 'QB', a deliberate A/B result. But that one flag also
silently decides which row-matcher runs:
- QB →
_find_player_rows— best candidate by row count, and explicitly no team+position fallback ("does NOT fall back to team+position soup"). - RB/WR/TE → the legacy cascade, which does fall through to that soup.
89 of 500 sampled players get different rows from the two paths. A back with
4 games of his own can end up with ~150 rows of team-position average as his
"player-specific" baseline — which then passes the >= 5 rows check. That looks
exactly like a leak the v4 design had already rejected.
Two separate concerns coupled through one flag, so the matcher was decoupled
behind unified_player_matching (default off) and measured.
Result: the clean matcher is worse, in both seasons
Walk-forwardwalk-forwardEvaluating week by week using only what was knowable before each week, mimicking how the model would actually have been used at the time. MAEmean absolute errorAverage size of the miss, ignoring direction. If a projection is off by 3 one week and -5 the next, the MAE is 4. Lower is better., weeks 5-18. Positive = unified matcher has higher error. QB categories are byte-identical, as expected — they already use the clean path.
| category | 2024 Δ | 2025 Δ | |
|---|---|---|---|
| rushing_yards | +0.209 | +0.564 | worse both |
| rushing_attempts | +0.049 | +0.149 | worse both |
| receiving_yards | +0.047 | +0.157 | worse both |
| receptions | +0.012 | +0.007 | worse both |
| targets | +0.013 | +0.012 | worse both |
| rushing_tds | +0.002 | +0.002 | worse both |
| receiving_tds | −0.001 | +0.001 | mixed |
6 of 7 worse in both seasons, 0 better in both. rushing_yards flips from
beating the naive baseline to losing to it in both years (2024 −0.110 →
+0.099; 2025 −0.203 → +0.361).
Why — and it validates the QB-only config
For a player with thin history the choice is between two priors:
- the team+position average (what the soup gives), or
- the position baseline (what the clean path falls back to).
The config comment on shrinkage_positions already says why the second is bad
for skill positions: "RB/WR/TE baselines mix WR1s with WR3s and drag good/bad
players toward a mid-tier mean." The soup is a crude but real team-context
prior — a backup on a run-heavy offence gets that offence's profile rather than
the league's. For QB it loses, because the starter-QB position baseline is
already informative and there is only one starter per team.
So the coupling is accidental, but the behaviour it produces is correct at both ends. That is worth knowing explicitly rather than by luck.
Shipped
Nothing changes by default. unified_player_matching stays False, and the
flag plus this document record that the soup was measured, not overlooked —
so the next person who spots it does not "fix" it and quietly lose 0.2-0.6
yards of rushing MAE.