New to these reports? Start here
  • Dotted-underlined words have a plain-English definition — hover or tap them. Every term is also on the glossary page.
  • "Null" means we found nothing, not that something broke. Most reports here are negative results, on purpose — knowing an idea doesn't work is the point.
  • Two questions get asked separately. First, is the effect real? Second, is it already priced into the betting odds? An effect can be completely real and still useless to bet on.
  • A "calibration" row is a self-check. It runs the same method on something already known to be true. If that fails, the whole report is unreliable — so it's reported alongside the findings.
  • If a confidence interval includes zero, the real effect might be nothing at all, so no claim gets made.

What is worth borrowing from Silver's ELWAY post

2026-09-08. Assessment of the free portion of natesilver.net/p/2026-elway-nfl-ratings-projections-playoff-odds, pasted by the user (the page is paywalled and renders as a JS shell from this server).

Method only. Silver Bulletin remains excluded as a data source on ToS grounds; nothing here ingests his ratings.

Most of it we already do

The post's framing — "most NFL ratings in the public domain are based on only a minimal amount of information, often just the final scores" — does not describe our model. features_temporal.py already carries EPAexpected points addedHow much a play changed the number of points its team should expect to score on that drive. A better measure of a play than raw yards., yards, turnovers, an EloElo ratingA rating system, originally from chess, that moves a team up or down based on results and the strength of the opponent. term, dynamic per-team home-field advantage, rest and fatigue, and injury status. So these ELWAY claims are not new information:

ELWAY feature us
component/efficiency stats over points already in (EPA, yards, turnovers)
per-team home-field advantage already in (dynamic HFA)
rest already in
injuries → point spread partly (practice status shipped into projections)
weather measured and rejected for betting — wind r² 1.1%, 50.0% unders at wind≥15
offseason roster WAR not done; see the player-movement TODO

Two of his ideas are things we have already measured and buried: weather has no betting edge here, and preseason signal is a dead nullnull resultA test that found nothing. "Null" is the starting assumption that there is no real effect; a "null result" means the data gave us no reason to abandon that assumption. It does not mean the data was missing or the test failed to run. (teams r=0.12, rookies r≈0). If the paid section claims a preseason edge, that is the claim to test hardest, not to adopt.

The one idea worth taking: realistic score simulation

"ELWAY uses a simulation engine to produce realistic game scores: 24-21 games occur much more often in ELWAY than 25-20."

JointScoreModel does the opposite. It turns a predicted margin into a win probability with a continuous Normal CDF, P(home) = Φ(margin/σ), and there is no key-number handling anywhere in the model.

Raw margins are nothing like Normal (6,758 games, 2001-2025):

margin = k empirical Normal ratio
3 0.1503 0.0529 2.84x
7 0.0906 0.0483 1.87x
14 0.0490 0.0345 1.42x
10 0.0550 0.0430 1.28x
5 0.0358 0.0510 0.70x
13 0.0272 0.0367 0.74x

⚠ But most of that structure disappears once you condition on the line, which is the case that actually applies to us. Residuals against an integer closing lineclosing lineThe final odds right before a game starts. It reflects everything the betting market knows, which makes it the hardest benchmark to beat. (n=3,754) are close to Normal — 3 at 0.90x, 7 at 1.07x. The line sits on the key numbers itself and absorbs them. Anyone reading only the table above would overstate the gain.

The defect that survives conditioning is the push. Residual = 0 occurs 1.56x more than Normal, and by line:

line n P(push)
exactly 3 1,083 0.091
exactly 7 444 0.056
exactly 10 199 0.065
any integer 3,754 0.047
half-point 3,004 0.000 ✓

Our Φ assigns ~0 to every one of those. On a 3-point line, nine games in a hundred push and the model prices that as impossible. The half-point row coming out at exactly 0.000 is the sanity check that the measurement is right.

Recommendation

  1. Give the model push mass at integer lines. This is the whole of the measurable defect and does not need a simulation engine — an empirical residual distribution would do it. Same fix pattern as the floor/ceiling band repair, which also replaced a parametric assumption with empirical quantiles.
  2. A score simulator is the fuller version and would also fix totals and give correct tie/overtime handling, but the marginal gain over (1) for spreads specifically is small, because the line already absorbs the lumpiness. Do not build it expecting a spread edge.
  3. Playoff odds are a genuine feature gap — we produce none for the NFL. That is a product decision, not an edge.
  4. ⚠ None of this is a betting edge. The market prices key numbers better than we do; that is why 3 and 7 are sticky. This fixes our own published numbers, which is worth doing on its own terms.

Built (2026-09-08)

Recommendation 1 shipped. JointScoreModel.cover_proba / over_proba now price off a discrete score distribution instead of Φ((m-line)/σ):

P(M = m)  ∝  φ((m - pred)/σ) · w(|m|)

w is an empirical key-number multiplier fitted on 2001-2018 and frozen in data/models/nfl_margin_keynumbers.json. The fitted shape is the known NFL structure, recovered from the data rather than asserted:

|margin| 1 2 3 4 5 6 7 8 10 14
w 0.73 0.71 2.91 0.97 0.65 1.19 1.95 0.74 1.34 1.44

Validated on 2019-2025 (n=1,960), three-way so the push is actually scored:

log losslog lossA score for probability forecasts that punishes confident wrong answers harshly. Lower is better. expected pushes actual
Gaussian (what shipped) 1.1326 0 43
key-number discrete 0.7614 43 43

⚠ Read that honestly: the cover probability itself did not improve. On non-push games the two-way log loss goes 0.6931 → 0.6927, +0.0004. The entire gain is mass the Gaussian could not express. Push calibrationcalibrationA deliberate sanity check on the method itself: run it on something already known to be true. If it fails to detect the known thing, the method is broken and its other results mean nothing. on held-out seasons is 0.047 predicted vs 0.047 actual across all integer lines (per-line n<50 is noisy: the 10 bucket predicts 0.039 against 0.121 on 33 games).

Totals are discretized but not key-number weighted. They push 2.9% of the time on an integer line, and a plain discrete distribution already reproduces that; the ratios (0.89-1.55x) show none of the sharp structure margins have, so weighting them would assert something unmeasured.

win_proba is deliberately untouched — the moneyline was never the defect.

Downstream, /models was scoring spread and total records with no push branch, so a push counted as a loss in the record, the ROIreturn on investmentProfit as a percentage of the money wagered. +2% means $2 profit per $100 bet. and the log loss. Pushes are now excluded from all three, which is both what a bettor experiences and what the model now predicts separately.

On this page

Terms in this report

Source

backtests/elway_review/FINDINGS.md
updated 2026-09-08 03:45