New to these reports? Start here
  • Dotted-underlined words have a plain-English definition — hover or tap them. Every term is also on the glossary page.
  • "Null" means we found nothing, not that something broke. Most reports here are negative results, on purpose — knowing an idea doesn't work is the point.
  • Two questions get asked separately. First, is the effect real? Second, is it already priced into the betting odds? An effect can be completely real and still useless to bet on.
  • A "calibration" row is a self-check. It runs the same method on something already known to be true. If that fails, the whole report is unreliable — so it's reported alongside the findings.
  • If a confidence interval includes zero, the real effect might be nothing at all, so no claim gets made.

NFL Edge Research — July 2026 (offseason build window)

Date: 2026-07-02. All code in this directory; nothing in production was touched. All results below are graded against closing lines, flat stakes, pushes excluded, one bet per pick (dedup'd). Given the July 2026 track-record audit, every historical "NFL record" found on disk was re-verified before use — two of them turned out to be retrodictions (see 1.3).


1. Current NFL edge surface (what we exploit today: effectively nothing)

1.1 Live pipeline. [redacted] contains zero NFL alerts ever (store began 2026-02-01, after the NFL season). [redacted] (source tag nfl_model_v3) is wired to scan h2h + spread + total vs devigged book lines using the temporal model + residual sigmas (sigma_spread~12.67, sigma_total~13), but it has never produced a live, settled record.

1.2 Model quality (honest, walk-forwardwalk-forwardEvaluating week by week using only what was knowable before each week, mimicking how the model would actually have been used at the time.). From experiments/iter_03_FINAL_holdout_2025.json (2025 holdoutholdoutData deliberately set aside and never looked at while developing an idea, then used once at the end as a fair test. Peeking at it first would defeat the purpose., leak-free): - ML Brier 0.2116 / accuracy 65.8% — vs the closing market's 0.2109 / 66.2% (measured in 04_h2h_market_efficiency.py). The model is at market level, not above it. - ml_flat betting at close: -5.26% ROIreturn on investmentProfit as a percentage of the money wagered. +2% means $2 profit per $100 bet. (n=179). The +6.5% "kellyKelly criterionA formula for how much to stake given your edge. Full Kelly maximizes long-run growth but swings wildly; most people use a fraction of it." ROI in the same file is on 2.46 total units staked — noise, not evidence. - Totals model backtest (data/totals_backtest_predictions.csv, 861 games 2021-2024): ROI -2.0% to -3.2% at every edge threshold (|edge|>=0/2/4/6). No edge in the totals model.

1.3 Invalid "records" (do not cite, ever). - Root model_performance.db: 1,087 NFL predictions all created in a 32-min window on 2026-01-18 covering 2021-2024 — hindsight retrodiction (already flagged by the audit). - [redacted]::stat_prop_bets claims +18.9% ROI on 4,431 dedup'd prop picks (62.4% win rate). Verified: this is a retrodiction written by scripts/backfill_stat_props.py on 2025-12-27 covering 2024 weeks 1-16 and 2025 weeks 1-17, with mismatched alt-line odds (e.g. an OVER graded at the standard 69.5 line but paid at +800 alt pricing; 25% "push" rate on X.5 lines). The number is fabricated-by-construction. Any props work must start from zero live evidence.

Bottom line: the platform currently has no proven NFL edge and no live NFL record.


2. Data inventory (what's actually on disk / freely reachable)

Source Coverage Usable for
nflverse nfl_data_py (internet OK; snapshot saved to schedules_1999_2025.feather) 1999-2025, 7,276 games: closing spread/total/moneylines, rest days, div flag, roof/temp/wind, QB/coach/ref teasers, situational, market-efficiency, weather
[redacted] 42,118 spread rows + 24,564 total rows, 2020-2024, 10-23 books/game, snapshot = Sunday 12:00Z of game week (~5h pre-kick for Sunday slate; no TNF, MNF a day early) multi-book dispersion, line-shopping
[redacted] 11,981 NFL prop rows, but only 20 games (2025-12-27 to 2026-02-01), 3 books, ~6 snapshots/prop, incl. Q1 variants too thin to backtest; proves the capture pipeline works
[redacted] (922 MB) Feb-Jun 2026 NBA/NHL/MLB only — no NFL nothing for NFL
[redacted] line_movement_history 27 games, Jan 2026 only nothing
Early-week -> close NFL line paths absent must capture live
Alt lines / teaser prices / derivatives (1H, team totals) absent must capture live

3. Candidate edges tested (scripts 01_-05_, reproducible)

3.1 Wong teasers (6-pt, 2-team) — 01_wong_teasers.pyBEST CANDIDATE

Rule fixed a priori (Wong 2001): tease dogs +1.5..+2.5 -> +7.5..+8.5, faves -7.5..-8.5 -> -1.5..-2.5 (both cross 3 and 7). Per-leg breakevens: 72.4% at -110, 73.9% at -120, 75.8% at -135 (current DK/FD standard).

Segment n legs cover 2-team ROI @-120 @-135
All legs 2015-2025 747 76.2% +6.4% +1.0%
All legs 2021-2025 386 75.1% +3.5% -1.7%
Dog legs 2015-2019 188 77.7% +10.6% +5.0%
Dog legs 2020-2025 303 77.6% +10.3% +4.7%
Fav legs 2020-2025 144 72.2% -4.4% -9.2%

Honest read: dog legs only are consistent across two independent 5-6 season eras (77.7% / 77.6%) and clear the -120 breakeven by ~2 SE; they clear the -135 breakeven by only ~1 SE. Fav legs have decayed below breakeven — the market fixed that half of the strategy. The dogs/favs split is post-hoc here but matches the published literature (favorite Wong legs died first). Volume is small: ~50 dog legs/season ~= 25 two-team teasers/season. Verdict: real but modest, and entirely price-dependent — +EV at <=-125, roughly breakeven-to-thin at -135. Requires shopping teaser pricing.

3.2 Wind unders — 02_wind_totals.py — NOT CONFIRMED

Threshold fit on 1999-2015 (wind >=18 mph: 61.1% under, +16.7% ROI), evaluated once on 2016-2025: 54.8% under, n=73 (SE +/-5.8%), +4.6% ROI — statistically indistinguishable from breakeven. The market demonstrably already prices wind (closing totals drop ~2.3 pts from calm to 21+ mph). ~7 qualifying games/season, and live betting relies on wind forecasts, which makes real results strictly worse than this backtest. Verdict: no deployable edge.

3.3 Situational ATS angles — 03_situational_ats.py — DEAD

12 classic angles x 2 eras vs closing spread. In 2015-2025 nothing clears breakeven meaningfully: post-bye -7.3%, divisional home dogs -4.5%, rest-edge

=4 days -3.0%, primetime dogs -3.2%. The nominal winner (freezing outdoor home teams, 56.4% n=140) is exactly the ~1 false positive expected from 24 tests. Verdict: fully priced; do not build.

3.4 h2h market efficiency / NHL-style stacking — 04_h2h_market_efficiency.py — NFL IS NBA-LIKE, NOT NHL-LIKE

Walk-forward logistic stack of [logit(devigdevigRemoving the bookmaker's cut from odds to recover the market's actual implied probability. closing ML), logit(MOV-Elo)], train 2007..T-1, test T for 2018-2025: - Stack Brier vs market Brier: mean delta +0.00001 (no improvement, any season). - EloElo ratingA rating system, originally from chess, that moves a team up or down based on results and the strength of the opponent. coefficient shrinks toward zero over time (+0.10 in 2018 -> +0.01-0.04 by 2024-25): the closing lineclosing lineThe final odds right before a game starts. It reflects everything the betting market knows, which makes it the hardest benchmark to beat. has absorbed everything public-data models know. - Betting sim (>=2% edge vs devig, paid at actual closing price): -9.0% ROI, n=207. - Corroborated by the platform's own stronger model: -5.3% flat at close (1.2). Verdict: do NOT attempt an NFL h2h/spread/total model-vs-close edge. The NHL v3 playbook does not transfer; NFL close is the sharpest line we track.

3.5 Off-consensus line shopping — 05_multibook_line_shopping.py — OVERLAY, NOT AN EDGE

2020-2024, 1,408 games, near-close multi-book snapshots. Bet any book >=0.5 pt better than the side's median consensus (dedup per game-side): - n=1,522: +0.7% ROI at the outlier line/price vs -1.8% for the same sides at consensus -110 -> line shopping is worth ~2.5% ROI. - But >=1.0 pt outliers lose -4.2% (n=337): big deviations are informed books being right (or limits/staleness you can't actually bet), not free money. Verdict: shopping recovers most of the vigvigThe bookmaker's built-in cut. It is why a coin-flip bet at -110 needs you to win about 52.4% of the time just to break even. but is ~breakeven standalone. Use as an execution overlay on any real edge (e.g., teaser pricing); never chase large outliers.

3.6 Player props / early-week CLV capture — INSUFFICIENT DATA

  • Genuine prop capture covers 20 games (late Dec -> SB). No honest backtest possible; the only "record" (+18.9%) is fabricated (1.3).
  • Zero early-week->close NFL line history exists on disk, so "capture CLVclosing line valueWhether you got a better price than the market settled at. Widely used as a faster signal of skill than profit, which takes ages to measure. as the edge" (bet Tue/Wed numbers that beat the close) is untestable today. Verdict: cannot be evaluated until we capture a season of data — see 4.

4. Ranked recommendation for the 2026 season

  1. Ship a Wong-teaser dog-leg scanner (only proven +EV item). Rule: current spread +1.5..+2.5, 6-pt tease, pair legs across games, only alert when 2-team pricing <= -125 (pass at -135 unless a book discounts). Honest expectation: low-single-digit % ROI on ~25 teasers/season — a small, real edge, not a headline. Track it live from week 1 with true timestamps so a real record exists by January.
  2. Start capture in August (this is the actual high-value move): (a) NFL spreads/totals/ML snapshots from Tuesday open through close (extend the existing unified_odds poller — it already does this for other sports); (b) full-season prop lines with open+close (extend the player_props poller beyond the 20-game trial); (c) teaser/alt-line pricing per book. This makes early-week-CLV, props, and derivatives honestly testable by mid-season 2026 / offseason 2027.
  3. Line-shop everything (~+2.5% execution overlay), but never treat an off-market line as signal — >=1 pt outliers lose.
  4. Explicitly do NOT build: NFL h2h/spread/total model-vs-close edges (market is NBA-efficient; our own model loses 5% at close), situational angles, wind unders. And never resurface stat_prop_bets or root model_performance.db numbers — both are retrodictions.

Reproduce

python3 01_wong_teasers.py          # teaser legs, eras, prices
python3 02_wind_totals.py           # train/test wind threshold
python3 03_situational_ats.py      # 12 angles x 2 eras
python3 04_h2h_market_efficiency.py # stacking test + betting sim
python3 05_multibook_line_shopping.py

Data snapshot: schedules_1999_2025.feather (nflverse, downloaded 2026-07-02).

On this page

Terms in this report

Source

backtests/nfl_edges/FINDINGS.md
updated 2026-07-02 13:14