New to these reports? Start here
- Dotted-underlined words have a plain-English definition — hover or tap them. Every term is also on the glossary page.
- "Null" means we found nothing, not that something broke. Most reports here are negative results, on purpose — knowing an idea doesn't work is the point.
- Two questions get asked separately. First, is the effect real? Second, is it already priced into the betting odds? An effect can be completely real and still useless to bet on.
- A "calibration" row is a self-check. It runs the same method on something already known to be true. If that fails, the whole report is unreliable — so it's reported alongside the findings.
- If a confidence interval includes zero, the real effect might be nothing at all, so no claim gets made.
Volume features: do they add anything on top of the snap re-ranker?
Date: 2026-08-30
Harness: backtests/fantasy_eval/volume_features_experiment.py (re-runnable)
Verdict: ✅ SHIPPED — small but consistent incremental gain
The question that actually mattered
The audit's most replicated result was volume is stable, efficiency is noise,
and nfl_stats.db holds 39,385 player-weeks of target_share /
air_yards_share that no model read. The obvious move is to feed them in.
But snap share already shipped as a start/sit re-ranker the same day, and snap share and target share measure overlapping things. So "does volume predict?" is the wrong question — it does, trivially. The question is whether volume adds anything on top of what is already deployed. Four nested models, identical walk-forwardwalk-forwardEvaluating week by week using only what was knowable before each week, mimicking how the model would actually have been used at the time., SpearmanSpearman correlationA measure of whether two rankings agree, from -1 (opposite) through 0 (unrelated) to +1 (identical). Cares about order, not exact values. within position-week, bootstrapbootstrapRe-running a calculation on thousands of resampled versions of the data to see how much the answer wobbles. The spread of those answers becomes the confidence interval. over cells:
| model | Spearman | vs base | 95% CI95% confidence intervalThe range the true value is plausibly in. If this range includes zero, we cannot rule out that the real effect is nothing at all. | MAEmean absolute errorAverage size of the miss, ignoring direction. If a projection is off by 3 one week and -5 the next, the MAE is 4. Lower is better. |
|---|---|---|---|---|
| base (projection only) | 0.4971 | — | — | 5.167 |
| + snap (what ships) | 0.5272 | +0.0301 | [+0.0187, +0.0467] | 5.052 |
| + volume | 0.5196 | +0.0226 | [+0.0136, +0.0368] | 5.121 |
| + both | 0.5339 | +0.0368 | [+0.0243, +0.0522] | 5.015 |
Incremental (+both − +snap) = +0.0067, 95% CI [+0.0032, +0.0103]. CI excludes zero, and it is positive in every test season: 2023 +0.0063, 2024 +0.0085, 2025 +0.0054.
Close-band start/sit (adjacent pairs within 2 projected points), which is the decision this surface exists to make:
| model | correct |
|---|---|
| base | 54.88% |
| + snap | 58.78% |
| + both | 59.44% |
Volume is worth about a fifth of what snaps are worth. Small, real, cheap — same infrastructure, one more table read.
⚠️ It only helps when the training set is big enough
Replicating 2025 with a within-season-only training set (the small-N regime) gives +0.0199, below the snap-only +0.0224 — the extra six columns spend degrees of freedom the data cannot pay for. The multi-season fit that production actually uses (~12,000 rows) is where the +0.0067 appears.
So the refit threshold now scales with the feature count
(MIN_ROWS_PER_FEATURE = 60 → 720 rows for the 12-feature model) instead of
staying at the 250 that was calibrated for the 6-feature snap model. Below that
floor the re-ranker declines and the order is unchanged.
⚠️ nflreadpy rejects numpy integers
load_snap_counts(seasons=[np.int64(2021)]) raises "Season must be between
2012 and 2025" for an in-range season. In rank_features that exception was
caught and turned into a silent no-op — the first run of this experiment
reported +snap as exactly identical to base, CI [0.0000, 0.0000], because
every snap feature was missing. _season_features now casts to int.
Any code passing a pandas/numpy season into nflreadpy has this bug.
Coefficients (standardized, fit on 12k rows before 2025 wk10)
proj +3.615 target_share_l3 +0.605
snap_l3 +1.267 carries_l3 +0.647
snap_trend +0.605 air_yards_share_l3 -0.196
snap_games +0.206 targets_l3 -0.035
pos_rank -0.257 receptions_l3 -0.113
Signs are physically sensible: more snaps, more target share and more carries all mean more points; a higher (worse) depth rank means fewer.
What this does NOT claim
Unchanged: the projection engine and its MAE — this is a re-ranking layer over the projections and writes no projection field. The MAE column above is the ridge's own fitted prediction, not the engine's output, and is not comparable to the engine's headline number.