Research Library
Every findings document in the repo, rendered. Sourced directly from the
markdown under backtests/ and docs/ — nothing to
register, a new file shows up here automatically. Negative results are kept
deliberately: most of what's here is a hypothesis that didn't survive.
37 documents are private and not shown.
All areas
backtests
bet_timing
cfb_eval
clv_line_adjustment
college_hockey_eval
docs
elway_review
ev_clv_prereg_2026_09
ev_record_2026_08
fantasy_eval
fantasy_model_audit
hr_derby
injury_impact
kelly_sim
kicker_dst
leg_correlation
middles
ncaab_eval
nfl_context_signals
nfl_edges
nfl_game_quality
nfl_prop_usage
nfl_war
nhl_eval
pickem
pickem_tiebreak
player_movement
preseason_signal
qb_handedness
qb_weather
split_miner
stable_filter_audit
stacking_eval
strategies
tennis_eval
weather_ftn_value
fantasy_eval · 16
Do the Injury-Aware Multipliers Actually Improve Projections?
Validation of commit 02fba14, which replaced the projection engine's injury multiplier table using results from backtests/nfl_context_signals/FINDINGS.md.
backtests/fantasy_eval/FINDINGS_INJURY_MULTIPLIERS.md
If an Injured Player Suits Up, What Does He Score?
shipped
Follow-up to FINDINGS_INJURY_MULTIPLIERS.md, which shipped an expected-value injury multiplier and left the conditional one — "if he is active, what should I expect?" — as an aside: ~0.82 / 0.92 / 0.97 relative to healthy for DNP / Limited / Full. Nothing in t…
backtests/fantasy_eval/FINDINGS_INJURY_CONDITIONAL.md
Depth-chart multiplier damp — re-measured after the input fix (2026-09-12)
Verdict: keep 0.4. Mean error is flat across the damp; what moves is bias.
backtests/fantasy_eval/FINDINGS_DEPTH_DAMP.md
Snap Counts + Depth Charts Close ~30% of the ECR Gap
Date: 2026-07-19 · Season: 2025, weeks 7–17 · Harness: snap_depth_experiment.py → snap_depth_experiment_results.json · Seed: 20260719
backtests/fantasy_eval/FINDINGS_SNAP_DEPTH.md
Cross-player usage redistribution: the effect is real, the projection gain is not
rejected
Date: 2026-08-31 · Verdict: ❌ REJECTED — usage_redistribution = False
backtests/fantasy_eval/FINDINGS_USAGE_REDISTRIBUTION.md
Backup QB: shade the receivers, not the quarterback
shipped
Date: 2026-08-31 · Verdict: ✅ SHIPPED, receiving only
backtests/fantasy_eval/FINDINGS_BACKUP_QB.md
Train / dev / holdout: what the accuracy numbers are actually worth
Date: 2026-08-30 · Report: scripts/report_split_accuracy.py (re-runnable)
backtests/fantasy_eval/FINDINGS_EVAL_SPLIT.md
Two dead knobs: one was worth fixing, one was worth measuring and leaving off
shipped
Date: 2026-08-30 · Harness: evaluate_fantasy_mae.py --set <knob>=true
backtests/fantasy_eval/FINDINGS_PACE_AND_ROSTERS.md
Start/sit accuracy of the shipped surface
shipped
Date: 2026-08-30 · Harness: start_sit_live_accuracy.py (re-runnable)
backtests/fantasy_eval/FINDINGS_START_SIT_LIVE.md
Floor/ceiling were covering half the outcomes they claimed
Date: 2026-08-30 Verdict: ✅ FIXED — empirical residual quantiles replace the simulated band
backtests/fantasy_eval/FINDINGS_INTERVALS.md
The team+position "soup" is not a bug — it beats the clean matcher
rejected
Date: 2026-08-30 Harness: scripts/evaluate_fantasy_mae.py --set unified_player_matching=true|false Verdict: ❌ REJECTED — keep unified_player_matching = False
backtests/fantasy_eval/FINDINGS_PLAYER_MATCHING.md
Volume features: do they add anything on top of the snap re-ranker?
shipped
Date: 2026-08-30 Harness: backtests/fantasy_eval/volume_features_experiment.py (re-runnable) Verdict: ✅ SHIPPED — small but consistent incremental gain
backtests/fantasy_eval/FINDINGS_VOLUME_FEATURES.md
TD models: not unpredictable, just pointed at the wrong output
Date: 2026-08-30 Harness: backtests/fantasy_eval/td_probability_experiment.py (re-runnable) Verdict: ✅ PARTIAL — ships for rushing and receiving TDs, NOT for passing TDs
backtests/fantasy_eval/FINDINGS_TD_PROBABILITY.md
Our Fantasy Projections vs FantasyPros Expert Consensus (ECR)
Our side is the strict walk-forward regeneration from scripts/evaluate_fantasy_mae.py --dump-projections (2,438 player-weeks): the engine is fed one historical week at a time and may only use data from before the week it projects. Unlike the 2024 db rows, this…
backtests/fantasy_eval/FINDINGS_VS_ECR.md
The Close Band Can Be Improved (+2.5pp, ~half the ECR gap)
Date: 2026-07-20 · Harness: close_band_experiment.py · Sample: 12,510 close pairs of 56,046, 2025 weeks 5–17 · Seed: 20260719
backtests/fantasy_eval/FINDINGS_CLOSE_BAND.md
Start/Sit Decision Accuracy — 2025
Date: 2026-07-19 · Harness: start_sit_accuracy.py · Sample: 64,150 within-(week, position) pairs, 2025 weeks 5–17 · Seed: 20260719
backtests/fantasy_eval/FINDINGS_START_SIT.md