New to these reports? Start here
- Dotted-underlined words have a plain-English definition — hover or tap them. Every term is also on the glossary page.
- "Null" means we found nothing, not that something broke. Most reports here are negative results, on purpose — knowing an idea doesn't work is the point.
- Two questions get asked separately. First, is the effect real? Second, is it already priced into the betting odds? An effect can be completely real and still useless to bet on.
- A "calibration" row is a self-check. It runs the same method on something already known to be true. If that fails, the whole report is unreliable — so it's reported alongside the findings.
- If a confidence interval includes zero, the real effect might be nothing at all, so no claim gets made.
Kicker and DST fantasy projections — feasibility study
2026-09-07. 3,614 kicker-games and 3,648 defense-games, 2019-2025 regular
season, built from nflverse play-by-play; closing lines from load_schedules
(100% coverage). Scoring: standard K (FG <40 = 3, 40-49 = 4, 50+ = 5, miss
= -1, XP = 1) and standard DST (points-allowed tiers, 1/sack, 2/turnover,
6/defensive TD, 2/safety).
The fantasy engine currently projects QB/RB/WR/TE only. K and DST are absent, so this asks whether they are worth adding, not whether they can be improved.
The proposed target is the wrong one
Field-goal attempts do not respond to game context. Across implied team total, FGA is flat:
| implied team total | <17 | 17-19 | 19-21 | 21-23 | 23-25 | 25-27 | 27+ |
|---|---|---|---|---|---|---|---|
| FGA | 1.76 | 1.90 | 2.03 | 2.01 | 2.02 | 1.99 | 1.88 |
| XPA | 1.53 | 1.65 | 1.96 | 2.26 | 2.51 | 2.89 | 3.27 |
| K pts | 6.59 | 7.16 | 7.65 | 7.75 | 8.16 | 8.59 | 8.69 |
corr(FGA, implied team total) = +0.017. The theorised inverted-U is there (bad offences don't reach range, great ones score touchdowns instead) but its amplitude is ~0.25 attempts across the entire range of NFL offences.
XPA is the part that scales — monotonemonotoneConsistently moving one direction as the input increases. A real dose-response effect should be monotone; if more of the cause does not mean more of the effect, the pattern is suspect. 1.53 → 3.27 — and XPA is just
"how many touchdowns will this team score", which JointScoreModel already
outputs. There is no separate model to build for it.
So a field-goal-attempt model targets the flat, noisy half of kicker scoring and skips the half that actually moves.
Kicker accuracy is noise, and attempt volume barely persists either
Season-to-season, same kicker, >=15 FGA in both years (n=155 pairs):
| r | |
|---|---|
| FG% | -0.026 |
| FGA/game | +0.135 |
| K points/game | +0.035 |
| team FGA/game (n=192) | +0.153 |
FG% has no persistence at all — the standard finding. Attempts explain 59% of within-game kicker points variance and realized accuracy would add 30 points more, but that 30 is unreachable: nothing predicts it.
Half of DST scoring is unpredictable by construction
Variance share of DST points, and how much of each piece Vegas explains (implied opponent total + game total + home):
| component | mean pts | var share | R² from Vegas |
|---|---|---|---|
| points-allowed tier | +0.49 | 0.326 | 0.134 |
| sacks | +2.40 | 0.174 | 0.039 |
| turnovers | +2.58 | 0.283 | 0.015 |
| defensive TDs | +0.75 | 0.211 | 0.004 |
Turnovers and defensive TDs are half the variance and essentially unforecastable. Points allowed is the only piece with real signal, and it is the piece a closing total already prices.
What actually beats what (out of sample, one season held out at a time)
Restricted to weeks with >=4 prior games in the season.
Kicker (n=1,814)
| predictor | R² | MAEmean absolute errorAverage size of the miss, ignoring direction. If a projection is off by 3 one week and -5 the next, the MAE is 4. Lower is better. | rank rho |
|---|---|---|---|
| league mean (constant) | 0.0000 | 3.46 | — |
| own prior-weeks average | -0.0955 | 3.63 | +0.056 |
| Vegas context model | +0.0187 | 3.40 | +0.162 |
Defense (n=2,368)
| predictor | R² | MAE | rank rho |
|---|---|---|---|
| league mean (constant) | 0.0000 | 4.42 | — |
| own prior-weeks average | -0.0689 | 4.55 | +0.134 |
| Vegas context model | +0.0781 | 4.28 | +0.298 |
⭐⭐ Season-to-date points — how everyone ranks kickers and defenses — is worse than a constant. Negative R² at every prior-game cutoff tested (K: -0.24 at >=1 game through -0.05 at >=12; DST: -0.16 through -0.02). It holds weak ordering information (rho +0.06 / +0.13) while being a bad point estimate, which is the same volume-predicts / efficiency-is-noise result this project keeps finding, in a position where the whole visible stat line is efficiency.
Recommendation
DST: worth building. Rank rho +0.298 out of sample from three inputs we already have, roughly double what season-to-date points gives, on a position that is streamed weekly — ranking is exactly the decision being made. Ceiling is low (R² 0.078) and must be stated on the page.
Kicker: not worth building as a model. rho +0.162, R² 0.019. If kickers are shown at all, show the implied team total and stop — anything more elaborate is fitting the noise. Do NOT build the field-goal-attempt model; FGA is the flat half.
Neither position should ever be ranked by season-to-date fantasy points.
Built (2026-09-08)
DST shipped as [redacted] + /fantasy/dst. Kicker not
built, per the recommendation above.
Two forms were evaluated walk-forwardwalk-forwardEvaluating week by week using only what was knowable before each week, mimicking how the model would actually have been used at the time. before anything was frozen:
| variant | R² | MAE | rho |
|---|---|---|---|
| A direct OLS on total DST points | 0.0722 | 4.29 | 0.288 |
| B components + tier integral | 0.0787 | 4.27 | 0.295 |
| league mean baseline | 0.0000 | 4.41 | — |
B wins on both metrics and is the explainable one, so the page can show projected points allowed, sacks and turnovers rather than one opaque number.
The tier term is an integral, not a lookup. Points allowed is projected as N(mu, sd) and the expectation is taken across the tier step function. Indexing the table instead would understate a 21-points-allowed projection by ~0.8 points, because it would score the whole distribution at the 0.0 tier. With sd -> 0 the integral reproduces the table exactly (guarded by a test).
Defensive touchdowns and safeties are set to the league mean, deliberately: R² 0.004 from the closing lineclosing lineThe final odds right before a game starts. It reflects everything the betting market knows, which makes it the hardest benchmark to beat. means there is nothing to model.
Coefficients are frozen in data/models/fantasy/dst_model.json alongside the
walk-forward metrics, and every projection carries them, so no page has to
guess how good the number is. A game with no posted line is skipped, never
defaulted.