New to these reports? Start here
  • Dotted-underlined words have a plain-English definition — hover or tap them. Every term is also on the glossary page.
  • "Null" means we found nothing, not that something broke. Most reports here are negative results, on purpose — knowing an idea doesn't work is the point.
  • Two questions get asked separately. First, is the effect real? Second, is it already priced into the betting odds? An effect can be completely real and still useless to bet on.
  • A "calibration" row is a self-check. It runs the same method on something already known to be true. If that fails, the whole report is unreliable — so it's reported alongside the findings.
  • If a confidence interval includes zero, the real effect might be nothing at all, so no claim gets made.

Kicker and DST fantasy projections — feasibility study

2026-09-07. 3,614 kicker-games and 3,648 defense-games, 2019-2025 regular season, built from nflverse play-by-play; closing lines from load_schedules (100% coverage). Scoring: standard K (FG <40 = 3, 40-49 = 4, 50+ = 5, miss = -1, XP = 1) and standard DST (points-allowed tiers, 1/sack, 2/turnover, 6/defensive TD, 2/safety).

The fantasy engine currently projects QB/RB/WR/TE only. K and DST are absent, so this asks whether they are worth adding, not whether they can be improved.

The proposed target is the wrong one

Field-goal attempts do not respond to game context. Across implied team total, FGA is flat:

implied team total <17 17-19 19-21 21-23 23-25 25-27 27+
FGA 1.76 1.90 2.03 2.01 2.02 1.99 1.88
XPA 1.53 1.65 1.96 2.26 2.51 2.89 3.27
K pts 6.59 7.16 7.65 7.75 8.16 8.59 8.69

corr(FGA, implied team total) = +0.017. The theorised inverted-U is there (bad offences don't reach range, great ones score touchdowns instead) but its amplitude is ~0.25 attempts across the entire range of NFL offences.

XPA is the part that scales — monotonemonotoneConsistently moving one direction as the input increases. A real dose-response effect should be monotone; if more of the cause does not mean more of the effect, the pattern is suspect. 1.53 → 3.27 — and XPA is just "how many touchdowns will this team score", which JointScoreModel already outputs. There is no separate model to build for it.

So a field-goal-attempt model targets the flat, noisy half of kicker scoring and skips the half that actually moves.

Kicker accuracy is noise, and attempt volume barely persists either

Season-to-season, same kicker, >=15 FGA in both years (n=155 pairs):

r
FG% -0.026
FGA/game +0.135
K points/game +0.035
team FGA/game (n=192) +0.153

FG% has no persistence at all — the standard finding. Attempts explain 59% of within-game kicker points variance and realized accuracy would add 30 points more, but that 30 is unreachable: nothing predicts it.

Half of DST scoring is unpredictable by construction

Variance share of DST points, and how much of each piece Vegas explains (implied opponent total + game total + home):

component mean pts var share R² from Vegas
points-allowed tier +0.49 0.326 0.134
sacks +2.40 0.174 0.039
turnovers +2.58 0.283 0.015
defensive TDs +0.75 0.211 0.004

Turnovers and defensive TDs are half the variance and essentially unforecastable. Points allowed is the only piece with real signal, and it is the piece a closing total already prices.

What actually beats what (out of sample, one season held out at a time)

Restricted to weeks with >=4 prior games in the season.

Kicker (n=1,814)

predictor R² MAEmean absolute errorAverage size of the miss, ignoring direction. If a projection is off by 3 one week and -5 the next, the MAE is 4. Lower is better. rank rho
league mean (constant) 0.0000 3.46 —
own prior-weeks average -0.0955 3.63 +0.056
Vegas context model +0.0187 3.40 +0.162

Defense (n=2,368)

predictor R² MAE rank rho
league mean (constant) 0.0000 4.42 —
own prior-weeks average -0.0689 4.55 +0.134
Vegas context model +0.0781 4.28 +0.298

⭐⭐ Season-to-date points — how everyone ranks kickers and defenses — is worse than a constant. Negative R² at every prior-game cutoff tested (K: -0.24 at >=1 game through -0.05 at >=12; DST: -0.16 through -0.02). It holds weak ordering information (rho +0.06 / +0.13) while being a bad point estimate, which is the same volume-predicts / efficiency-is-noise result this project keeps finding, in a position where the whole visible stat line is efficiency.

Recommendation

DST: worth building. Rank rho +0.298 out of sample from three inputs we already have, roughly double what season-to-date points gives, on a position that is streamed weekly — ranking is exactly the decision being made. Ceiling is low (R² 0.078) and must be stated on the page.

Kicker: not worth building as a model. rho +0.162, R² 0.019. If kickers are shown at all, show the implied team total and stop — anything more elaborate is fitting the noise. Do NOT build the field-goal-attempt model; FGA is the flat half.

Neither position should ever be ranked by season-to-date fantasy points.


Built (2026-09-08)

DST shipped as [redacted] + /fantasy/dst. Kicker not built, per the recommendation above.

Two forms were evaluated walk-forwardwalk-forwardEvaluating week by week using only what was knowable before each week, mimicking how the model would actually have been used at the time. before anything was frozen:

variant R² MAE rho
A direct OLS on total DST points 0.0722 4.29 0.288
B components + tier integral 0.0787 4.27 0.295
league mean baseline 0.0000 4.41 —

B wins on both metrics and is the explainable one, so the page can show projected points allowed, sacks and turnovers rather than one opaque number.

The tier term is an integral, not a lookup. Points allowed is projected as N(mu, sd) and the expectation is taken across the tier step function. Indexing the table instead would understate a 21-points-allowed projection by ~0.8 points, because it would score the whole distribution at the 0.0 tier. With sd -> 0 the integral reproduces the table exactly (guarded by a test).

Defensive touchdowns and safeties are set to the league mean, deliberately: R² 0.004 from the closing lineclosing lineThe final odds right before a game starts. It reflects everything the betting market knows, which makes it the hardest benchmark to beat. means there is nothing to model.

Coefficients are frozen in data/models/fantasy/dst_model.json alongside the walk-forward metrics, and every projection carries them, so no page has to guess how good the number is. A game with no posted line is skipped, never defaulted.

On this page

Terms in this report

Source

backtests/kicker_dst/FINDINGS.md
updated 2026-09-08 00:44