New to these reports? Start here
  • Dotted-underlined words have a plain-English definition — hover or tap them. Every term is also on the glossary page.
  • "Null" means we found nothing, not that something broke. Most reports here are negative results, on purpose — knowing an idea doesn't work is the point.
  • Two questions get asked separately. First, is the effect real? Second, is it already priced into the betting odds? An effect can be completely real and still useless to bet on.
  • A "calibration" row is a self-check. It runs the same method on something already known to be true. If that fails, the whole report is unreliable — so it's reported alongside the findings.
  • If a confidence interval includes zero, the real effect might be nothing at all, so no claim gets made.

Can we build our own WAR? What the data will and won't support

2026-09-08. 227,398 scrimmage plays, 2019-2025.

The /nfl/injuries leaderboard weights absences by snap share and says plainly that it is not WAR. This asks what it would take to make a real one.

The ceiling: 27% of the roster

Of 47,274 plays in 2024, 33,471 (71%) are scrimmage plays, and every one of those names an offensive player. But across the season only 599 distinct players are ever named, against 2,227 active players on rosters.

Nobody on the offensive line, defensive line, at linebacker, or in the secondary appears in play attribution at all. A play-by-play WAR can cover 27% of the league and no more. Any metric that claims to rate a guard or a safety from this data is inventing it.

Two honest options for the other 73%: leave them out and say so, or measure them a completely different way — with-or-without-you on the availability panel already built for backtests/injury_impact/, clearly labelled as a different thing.

What persists — and a trap that nearly went in the report

Year-over-year correlationcorrelationHow closely two things move together, from -1 (opposite) through 0 (unrelated) to +1 (in lockstep). It does not by itself mean one causes the other. of per-player-season measures:

role min plays volume (plays) EPAexpected points addedHow much a play changed the number of points its team should expect to score on that drive. A better measure of a play than raw yards. total EPA/play
passer 200 +0.361 +0.428 +0.475
rusher 80 +0.410 +0.443 +0.626
receiver 40 +0.586 +0.445 +0.351

⚠⚠ The rusher row is an artifact. The rusher pool contains quarterbacks, whose scrambles carry very different EPA from a running back's carries — and quarterbacks stay quarterbacks, so the population is bimodal and the correlation measures that, not skill. Excluding QBs:

rusher pool YoY r (EPA/play)
all rushers +0.626
QBs removed +0.096

Running-back rushing efficiency does not persist. That matches what the fantasy work found repeatedly and would have been reported backwards without the check.

How much of "skill" is really the situation

Splitting the same correlation by whether the player changed teams:

role stayed changed team
passer +0.464 (n=125) +0.318 (n=30)
rusher (no QBs) +0.085 (n=166) +0.142 (n=40)
receiver +0.365 (n=537) +0.199 (n=115)

Efficiency loses roughly a third of its persistence for quarterbacks and nearly half for receivers when the situation changes. What survives a team change is the part that is actually the player. A WAR that ignores this credits the offensive line and the scheme to whoever touched the ball.

The design the evidence supports

position what the data allows
QB a real efficiency-based WAR — persistence survives a team change
WR/TE usable, but efficiency shrunk hard toward the mean (only ~+0.20 of it is portable)
RB volume and receiving only. Rushing efficiency is +0.096 — do not build on it
OL/DL/LB/DB not from play-by-play. WOWY or nothing

Prototype: quarterback WAR, and it validates

Everything derived, nothing assumed:

  • points per win: 35.7, fitted from team point differential vs wins across 224 team-seasons (an external sanity check — the accepted NFL figure is ~35)
  • replacement level: −0.1286 EPA/dropback, the pooled rate of QBs outside the top 32 by volume in their season (555 QB-seasons, 23,159 dropbacks), against a starter average of +0.0478
  • WAR = (EPA/db − replacement) × dropbacks / 35.7

Top seasons, as a face-validity check:

player season dropbacks EPA/db WAR
P. Mahomes 2022 677 +0.258 7.3
P. Mahomes 2020 612 +0.275 6.9
A. Rodgers 2020 543 +0.325 6.9
J. Allen 2020 599 +0.266 6.6
L. Jackson 2024 494 +0.341 6.5

Out of sample, one season held out at a time, predicting team wins:

MAEmean absolute errorAverage size of the miss, ignoring direction. If a projection is off by 3 one week and -5 the next, the MAE is 4. Lower is better. (wins)
team QB WAR 1.840
starter dropbacks only 2.671
both 1.857
league mean (constant) 2.680

QB WAR cuts the error by 31% against a constant, and dropbacks add nothing on top of it — the metric already contains its own volume.

⚠ This is contemporaneous, not predictive: a season's QB WAR against that same season's wins. It shows the metric is scaled to wins correctly, which is what a WAR must be. It does not show it forecasts anything, and it should never be quoted as if it did.

Recommendation

Build QB WAR — it is the position the evidence supports, it validates, and per Silver's own account quarterbacks are about a third of player value. Add WR/TE with heavy shrinkage second. Do not build RB rushing WAR. Leave the trenches out of any play-by-play WAR, and if they are wanted, measure them by WOWY and call it something else.


Shipped (2026-09-08)

QB WAR is live. scripts/nfl/train_qb_war.py fits it and freezes it to data/models/nfl_qb_war.json; [redacted] serves it; /nfl/injuries shows an absent quarterback's WAR in its own column.

Wins and snap-shares are kept in separate columns on purpose. Adding them would be nonsense, and the whole reason to have QB WAR is that losing a starting quarterback is not the same as losing three special-teamers with the same combined snap share — which is exactly what a snap-share total says.

⚠⚠ A dropback floor was needed, and finding out why was the useful part. The first fit published a WAR for anyone who ever took a dropback. 350 of 779 QB-seasons had under 25 dropbacks and 134 of those carried a positive WAR: Courtland Sutton, a wide receiver, was credited +0.22 wins for two trick-play dropbacks at +3.745 EPA/db — 3% of a Brady MVP season. It surfaced on the leaderboard as Gunner Olszewski, a returner, being the only "quarterback" any team was missing.

The floor is 50 dropbacks, and anyone under it gets no WAR rather than zero — we cannot measure them, which is not the same as saying they are replacement level. Applying it removes 411 entries and costs nothing: out-of-sampleout-of-sampleTested on data that was not used to build or tune the idea. This is the honest test; results on the data you built with are almost always flattering. MAE moves 1.840 → 1.841 wins.


Is WAR contemporaneous by nature?

2026-09-09. Asked after the validation above was flagged as contemporaneous.

By design, yes. WAR is an accounting metric: it prices what a player actually did, in wins, over a period that already happened — and "what he actually did" includes his luck. Baseball's WAR is the same; projections (ZiPS, Steamer, PECOTA) are separate models built on top of WAR-like components, not WAR itself carried forward.

But "descriptive by design" is not "predictively empty", and the difference is measurable on our own metric.

A quarterback's WAR persists. Over 219 QB season-pairs with 50+ dropbacks in both:

year-over-year r
WAR +0.527
EPA/dropback +0.475
dropbacks +0.496

But most of the win-predicting power is contemporaneous. Predicting a team's wins, out of sample, one season held out at a time:

MAE (wins) corr
same-season team QB WAR 1.841 +0.690
prior-season team QB WAR 2.424 +0.366
league mean (constant) 2.674 —

Prior-season WAR keeps only about 30% of the improvement the same-season figure buys ((2.674−2.424) / (2.674−1.841)). It still beats a constant, so it is not worthless as a forecast — it is just a much weaker one.

How much carries, in win units:

wins per unit of team QB WAR
same season 1.076
prior season 0.600

⇒ about 56% of last year's QB WAR carries forward. The other 44% was opponent, scheme, health and luck. That ratio is the shrinkage a projection would have to apply, and it is what separates a WAR from a forecast.

⚠ Averaging more prior seasons did not help here — one season 2.424, two 2.582, three 2.552 — but the row counts fall as more history is required (192 / 160 / 128), so those are not like-for-like and this should not be read as evidence against multi-year priors.

So: quote WAR as what a player was worth, never as what he will be worth. To make it a projection you would regress it ~56% toward replacement and add an aging curve — at which point it is a different metric that should carry a different name.

On this page

Terms in this report

Source

backtests/nfl_war/FINDINGS.md
updated 2026-09-09 00:22