New to these reports? Start here
  • Dotted-underlined words have a plain-English definition — hover or tap them. Every term is also on the glossary page.
  • "Null" means we found nothing, not that something broke. Most reports here are negative results, on purpose — knowing an idea doesn't work is the point.
  • Two questions get asked separately. First, is the effect real? Second, is it already priced into the betting odds? An effect can be completely real and still useless to bet on.
  • A "calibration" row is a self-check. It runs the same method on something already known to be true. If that fails, the whole report is unreliable — so it's reported alongside the findings.
  • If a confidence interval includes zero, the real effect might be nothing at all, so no claim gets made.

The team+position "soup" is not a bug — it beats the clean matcher

Date: 2026-08-30 Harness: scripts/evaluate_fantasy_mae.py --set unified_player_matching=true|false Verdict: ❌ REJECTED — keep unified_player_matching = False

What prompted it

shrinkage_positions is 'QB', a deliberate A/B result. But that one flag also silently decides which row-matcher runs:

  • QB → _find_player_rows — best candidate by row count, and explicitly no team+position fallback ("does NOT fall back to team+position soup").
  • RB/WR/TE → the legacy cascade, which does fall through to that soup.

89 of 500 sampled players get different rows from the two paths. A back with 4 games of his own can end up with ~150 rows of team-position average as his "player-specific" baseline — which then passes the >= 5 rows check. That looks exactly like a leak the v4 design had already rejected.

Two separate concerns coupled through one flag, so the matcher was decoupled behind unified_player_matching (default off) and measured.

Result: the clean matcher is worse, in both seasons

Walk-forwardwalk-forwardEvaluating week by week using only what was knowable before each week, mimicking how the model would actually have been used at the time. MAEmean absolute errorAverage size of the miss, ignoring direction. If a projection is off by 3 one week and -5 the next, the MAE is 4. Lower is better., weeks 5-18. Positive = unified matcher has higher error. QB categories are byte-identical, as expected — they already use the clean path.

category 2024 Δ 2025 Δ
rushing_yards +0.209 +0.564 worse both
rushing_attempts +0.049 +0.149 worse both
receiving_yards +0.047 +0.157 worse both
receptions +0.012 +0.007 worse both
targets +0.013 +0.012 worse both
rushing_tds +0.002 +0.002 worse both
receiving_tds −0.001 +0.001 mixed

6 of 7 worse in both seasons, 0 better in both. rushing_yards flips from beating the naive baseline to losing to it in both years (2024 −0.110 → +0.099; 2025 −0.203 → +0.361).

Why — and it validates the QB-only config

For a player with thin history the choice is between two priors:

  • the team+position average (what the soup gives), or
  • the position baseline (what the clean path falls back to).

The config comment on shrinkage_positions already says why the second is bad for skill positions: "RB/WR/TE baselines mix WR1s with WR3s and drag good/bad players toward a mid-tier mean." The soup is a crude but real team-context prior — a backup on a run-heavy offence gets that offence's profile rather than the league's. For QB it loses, because the starter-QB position baseline is already informative and there is only one starter per team.

So the coupling is accidental, but the behaviour it produces is correct at both ends. That is worth knowing explicitly rather than by luck.

Shipped

Nothing changes by default. unified_player_matching stays False, and the flag plus this document record that the soup was measured, not overlooked — so the next person who spots it does not "fix" it and quietly lose 0.2-0.6 yards of rushing MAE.