New to these reports? Start here
- Dotted-underlined words have a plain-English definition — hover or tap them. Every term is also on the glossary page.
- "Null" means we found nothing, not that something broke. Most reports here are negative results, on purpose — knowing an idea doesn't work is the point.
- Two questions get asked separately. First, is the effect real? Second, is it already priced into the betting odds? An effect can be completely real and still useless to bet on.
- A "calibration" row is a self-check. It runs the same method on something already known to be true. If that fails, the whole report is unreliable — so it's reported alongside the findings.
- If a confidence interval includes zero, the real effect might be nothing at all, so no claim gets made.
If an Injured Player Suits Up, What Does He Score?
Follow-up to FINDINGS_INJURY_MULTIPLIERS.md, which shipped an
expected-value injury multiplier and left the conditional one — "if he
is active, what should I expect?" — as an aside: ~0.82 / 0.92 / 0.97 relative
to healthy for DNP / Limited / Full. Nothing in the repo measured those
numbers. This does.
Answer
| Questionable, practiced… | Conditional multiplier | 95% CI95% confidence intervalThe range the true value is plausibly in. If this range includes zero, we cannot rule out that the real effect is nothing at all. | n (played) | P(play) |
|---|---|---|---|---|
| any (pooled) | 0.87 | 0.84 – 0.90 | 1,521 | 0.64 |
| DNP | 0.76 | 0.67 – 0.85 | 146 | 0.39 |
| Limited | 0.87 | 0.83 – 0.92 | 1,048 | 0.66 |
| Full | 0.93 | 0.85 – 1.00 | 319 | 0.79 |
Relative to players not on the injury report, derived on 2015–2021, player-clustered bootstrapbootstrapRe-running a calculation on thousands of resampled versions of the data to see how much the answer wobbles. The spread of those answers becomes the confidence interval. (1,000 resamples, seeded and order-independent, so the intervals reproduce exactly). A Questionable player who plays scores about 13% below what his recent form implies; one who didn't practise all week, about 24% below.
Three things matter more than the table:
- These remove the biasbiasWhether the misses lean consistently one way. A projection can have a good average error size but still be biased if it is almost always too high. Bias is often the more fixable problem., not the error. On the 2022–25 holdoutholdoutData deliberately set aside and never looked at while developing an idea, then used once at the end as a fair test. Peeking at it first would defeat the purpose., conditional on playing, they take bias from −1.97 to −0.07 PPRpoints per receptionA fantasy scoring format that awards a point for every catch, which raises the value of high-volume receivers. — but MAEmean absolute errorAverage size of the miss, ignoring direction. If a projection is off by 3 one week and -5 the next, the MAE is 4. Lower is better. does not move (−0.012, CI −0.152 to +0.141). See Why MAE doesn't improve.
- The aside's numbers were measured against the wrong "healthy". Against players merely listed without a game designation they come out 0.80 / 0.92 / 0.98 — close to the aside. Against players not on the report at all, which is what "healthy" means to a manager, they are 0.76 / 0.87 / 0.93.
- No per-position table. The position pattern on 2015–21 flips on 2022–25.
Method
- Population and baseline identical to the EV study so the two are directly comparable: skill positions, trailing-4-game mean from strictly prior weeks (≥2 games), final pre-game injury report row per player-week.
- Step 0 reproduces the EV study first — conditional frame n = 818, MAE 5.067, bias −2.152, all exact — and the script refuses to continue if it doesn't. One discrepancy recorded rather than hidden: the holdout now has 2,504 designated player-weeks, not 2,503. Every conditional-frame number is unchanged, so the extra row is a non-player: an upstream nflverse injury-report correction since August, not a harness difference.
- Multiplier for class c = R_c / R_ref, where R = Σ actual PPR / Σ baseline over rows where the player played. Ratio of means, so a few near-zero baselines cannot dominate as they would in a mean of per-player ratios.
- Only played rows enter either reference, so bye weeks, cuts and healthy scratches — which make "not on the report" a messy class for availability — cannot touch these ratios.
- Out-of-sampleout-of-sampleTested on data that was not used to build or tune the idea. This is the honest test; results on the data you built with are almost always flattering.: derived on 2015–2021, evaluated on 2022–2025.
- Scored raw (prediction = baseline × multiplier, the EV study's convention) and calibrated (baseline × R_ref × multiplier, which strips out the trailing mean's own bias). They agree on every conclusion; raw is quoted below.
The reference class is not a detail
| Reference | R_ref (actual ÷ trailing-4) | n |
|---|---|---|
| Not on the injury report | 1.016 | 23,706 |
| Listed, no game designation | 0.964 | 3,605 |
The two differ by ~5%, and every multiplier inherits that. Players who show up on a practice report without a designation are themselves slightly below form — they're on the report for a reason.
This also corrects the record on the EV study: its docstring says it normalises against the not-on-report class, but the code uses listed, no designation.
The two tables are one measurement
With the EV study's own reference, EV = P(play) × conditional ÷ P(play | ref), and it reproduces the derived EV table to three decimals:
| P(play) | × conditional | ÷ P_ref | = implied EV | derived EV | |
|---|---|---|---|---|---|
| Questionable | 0.641 | 0.918 | 0.889 | 0.662 | 0.662 |
| Q-DNP | 0.386 | 0.796 | 0.889 | 0.346 | 0.346 |
| Q-Limited | 0.665 | 0.921 | 0.889 | 0.689 | 0.689 |
| Q-Full | 0.790 | 0.977 | 0.889 | 0.867 | 0.867 |
So no new data is needed to expose both — and it surfaces something about the shipped EV table, below.
Holdout 2022–25, conditional on playing (n = 818)
| Variant | MAE | Bias |
|---|---|---|
| No multiplier | 5.480 | +1.205 |
| Shipped EV, status only | 5.054 | −1.974 |
| Shipped EV + practice | 5.160 | −1.884 |
| Conditional, status only | 5.148 | −0.069 |
| Conditional + practice | 5.140 | −0.085 |
| Comparison | ΔMAE | 95% CI |
|---|---|---|
| Conditional vs no multiplier | −0.332 | −0.427 to −0.239 |
| Conditional (status) vs shipped EV + practice | −0.012 | −0.152 to +0.141 |
| Conditional + practice vs shipped EV + practice | −0.020 | −0.150 to +0.126 |
| Practice split vs status only, within conditional | −0.008 | −0.042 to +0.025 |
Injured players who play genuinely underperform — applying the conditional multiplier beats no multiplier clearly. But against the shipped EV table, the improvement is all bias and no MAE.
Why MAE doesn't improve
MAE is minimised by the median; bias measures distance from the mean. Fantasy outcomes for an injured player are right-skewed — a lot of low games, occasional big ones — so the median sits below the mean. The EV multiplier, biased two points low, lands near that median and ties the unbiased conditional multiplier on MAE.
Which one is right depends on the question:
- Summing points — a lineup total, a projected team score — needs the mean, and the EV table is two points low per active injured player there. The conditional multiplier is correct.
- Ranking one player against another — a start/sit call — is closer to a median comparison, and either works about equally.
So expose it as an "if active" figure next to the EV projection, not as a replacement — and don't sell it as an accuracy gain. It is a calibrationcalibrationA deliberate sanity check on the method itself: run it on something already known to be true. If it fails to detect the known thing, the method is broken and its other results mean nothing. fix.
The practice split: real ordering, no accuracy gain
DNP < Limited < Full on develop (0.76 / 0.87 / 0.93) and again on the holdout (0.81 / 0.86 / 0.90), and every holdout value sits inside its develop CI. The ordering is real. It buys nothing in MAE (−0.008, CI spans zero), mostly because DNP-but-played is rare: 146 player-weeks in seven seasons, 97 in four. Same shape as in the EV study — the split informs play probability far more than it informs production.
Position: suggestive in-sample, gone out of sample
| QB | RB | WR | TE | |
|---|---|---|---|---|
| Develop 2015–21 | 0.99 [0.90, 1.10] | 0.84 [0.77, 0.90] | 0.85 [0.80, 0.90] | 0.92 [0.84, 1.00] |
| Holdout 2022–25 | 0.86 [0.73, 1.08] | 0.92 [0.83, 1.01] | 0.85 [0.79, 0.91] | 0.81 [0.70, 0.94] |
"Questionable QBs who play are fine" is the story develop tells, and the holdout doesn't back it: QB drops to 0.86 and RB and TE swap places. Only WR holds. A per-position table would be fitting noise; the pooled or practice-split figures are the defensible ones.
RESOLVED 2026-09-26: the shipped EV table was normalised against a class that sits 11% of the time
The reconciliation shows the EV table divides by P(play | listed, no designation) = 0.889. The engine then applies it to a healthy projection whose own play probability it states as 1.0 (projections.py sets 'Healthy' to 1.0 deliberately, because for a startable player it effectively is). If the healthy projection assumes the player plays, the EV multipliers are inflated by roughly 1 / 0.889 for exactly the players anyone starts.
The EV study's holdout is consistent with that — EV-frame bias +0.29 overall and +0.54 on Questionable, i.e. still slightly over-projecting — but that study didn't isolate it, and neither does this one. Not changed here. A clean test is to re-derive the EV table with P_ref = 1 and the not-on-report reference, then re-score the EV frame on 2022–25 restricted to players with a fantasy-relevant baseline.
Tested and fixed — injury_ev_normalization.py. Derived on 2015–21, scored
on 2022–25 in the expected-value frame; the new table is
P(play) × production-if-he-plays ÷ production of a healthy player who played,
with no division by 0.889:
| Holdout 2022–25, with practice split | Old MAE / bias | New MAE / bias | ΔMAE, 95% CI |
|---|---|---|---|
| All designated (2,504) | 2.633 / +0.29 | 2.531 / −0.20 | −0.102 [−0.134, −0.069] |
| Questionable (1,322) | 4.972 / +0.54 | 4.779 / −0.38 | −0.193 [−0.254, −0.133] |
| Questionable, trailing-4 ≥ 8 (679) | 6.700 / +1.11 | 6.465 / −0.30 | −0.235 [−0.350, −0.127] |
Raw baseline shown; the calibrated baseline agrees (e.g. −0.109 [−0.142, −0.077] on all designated). The inflation was worst for startable players. Dropping only the 0.889 but keeping the listed reference sits in between (bias closest to zero overall, MAE slightly worse than new, and +0.27 bias on startable Questionables under the calibrated baseline), so the reference class matters too.
Shipped (pooled 2015–25, same formula): Questionable 0.55, DNP 0.32,
Limited 0.56, Full 0.71 (were 0.68 / 0.40 / 0.69 / 0.87). The old values
broke a simple invariant — a multiplier above the play probability means an
injured player outscores a healthy one whenever he plays — three times.
test_injury_projections.py now pins it.
Caveats
- The trailing-4 baseline for a multi-week injury may already include games played hurt, which pulls these ratios toward 1. The true effect of playing injured is, if anything, somewhat larger than measured.
- "Played" means recorded a stat line. A player who re-aggravates the injury in the first quarter counts as played, which is correct for this question.
- Doubtful (2 player-weeks) and Out (1) who played are too rare to estimate and are excluded from the table. Probable is 2015 only; the NFL abolished it.
Reproduce: python3 backtests/fantasy_eval/injury_conditional_ratio.py
(results in injury_conditional_ratio_results.json).