New to these reports? Start here
  • Dotted-underlined words have a plain-English definition — hover or tap them. Every term is also on the glossary page.
  • "Null" means we found nothing, not that something broke. Most reports here are negative results, on purpose — knowing an idea doesn't work is the point.
  • Two questions get asked separately. First, is the effect real? Second, is it already priced into the betting odds? An effect can be completely real and still useless to bet on.
  • A "calibration" row is a self-check. It runs the same method on something already known to be true. If that fails, the whole report is unreliable — so it's reported alongside the findings.
  • If a confidence interval includes zero, the real effect might be nothing at all, so no claim gets made.

Game quality — opponent-adjusted EPA, 2016–2025

2026-09-14. validate.py → validate_results.json. Shipped as /nfl/game-quality (scripts/nfl/build_game_quality.py → [redacted]).

The question: the final score hides how a game was won. "59 points" leaves out the 37 allowed; a one-point win leaves out a blown 21-point lead; a close loss leaves out a strip-sack on the last drive. Grade every play by EPAexpected points addedHow much a play changed the number of points its team should expect to score on that drive. A better measure of a play than raw yards. instead, then adjust for the opponent, so players and teams can be compared across different schedules.

Method

  • Per game: EPA per scrimmage play on offense and defense, and QB EPA per dropback (passes, sacks, scrambles). Turnover and sack EPA are split out, plus win probability added (WPA).
  • Opponent rating: leave-one-out. A defense's rating for game g is its per-game average over its other games that season, so the game being adjusted never rates its own opponent. It is shrunk toward a prior:

rating = (sum over other games + k × prior) / (n_other + k)

prior = league mean + w × (last season's team mean − last season's league mean)

  • Adjustment: adj_off = off − (opp_def_rating − league), and the same for defense and for QB EPA per dropback against pass defense.
  • Win expectancy: logistic P(win) on EPA margin, no intercept.

1. Shrinkage

Choose k and w by how well the rating predicts the held-out game's value, with each season cut off at week 2, 4, 8 and 18 so early-season behaviour counts equally.

Unit Best k Best w MSE No shrinkage
Defense EPA/play allowed 16 0.5 0.0402 0.0539
Offense EPA/play 12 0.5 0.0385 0.0508
Pass defense EPA/dropback 24 0.5 0.0881 0.1193
  • Shrinkage cuts error by about 25% for every unit. The first grid stopped at 12 and the best value sat on its edge, so the grid was widened to 48.
  • The curve is flat. k = 12 vs 16 differ in the fifth decimal. Shipped k = 16 for teams and k = 24 for pass defense, which is noisier per game.
  • Last season's gap is worth about half. w = 0.5 won in every unit.

2. Does the adjustment help? Odd vs even weeks

If the adjustment removes schedule noise, a team's odd-week and even-week numbers should agree more.

Measure Raw r Adjusted r n
Offense EPA/play 0.591 0.615 320 team-seasons
Defense EPA/play 0.316 0.362 320
EPA margin 0.523 0.551 320
QB EPA/dropback (≥75 dropbacks per half) 0.469 0.486 374 QB-seasons

The adjustment improves all four. The gains are modest but consistent, and largest for defense, the least stable unit.

3. Prediction: weeks 1–9 → weeks 10–18

Opponent ratings here use weeks 1–9 only, so nothing leaks from the future.

Raw Adjusted
Per-game point margin (2,626 games) r 0.354 r 0.359
Season point margin per game r 0.514 r 0.526

Same direction, small. The adjustment makes a rating slightly more predictive; it does not turn a description into a forecast.

4. Win expectancy

Coefficient 14.24 on EPA margin (5,502 team-games). Brier score 0.094, and it names the winner 86.5% of the time. So about 1 game in 7 is won by the team that lost the play-by-play battle. That is the "fortunate win" the page flags.

Caveats

  • Descriptive. One game's grade is what happened, not a rating of the team. The adjusted margin is only somewhat more stable than raw (above).
  • Quarterbacks are charged for sacks. Play-by-play puts every sack and dropback on the passer, so a bad offensive line shows up in the QB's numbers; sack EPA is a separate column so you can see how much.
  • Early-season adjustments lean on last year. With k = 16, a week-2 rating is about 90% prior.
  • Garbage time counts. Plays in decided games are included; the tested adjustment does not filter them.
  • The win-expectancy fit includes the seasons it scores (2 parameters on 5,500 games; negligible). The 2026 page uses a fit on completed seasons only.

On this page

Terms in this report

Source

backtests/nfl_game_quality/FINDINGS.md
updated 2026-09-14 01:23