Research Library
Every findings document in the repo, rendered. Sourced directly from the
markdown under backtests/ and docs/ — nothing to
register, a new file shows up here automatically. Negative results are kept
deliberately: most of what's here is a hypothesis that didn't survive.
37 documents are private and not shown.
All areas
backtests
bet_timing
cfb_eval
clv_line_adjustment
college_hockey_eval
docs
elway_review
ev_clv_prereg_2026_09
ev_record_2026_08
fantasy_eval
fantasy_model_audit
hr_derby
injury_impact
kelly_sim
kicker_dst
leg_correlation
middles
ncaab_eval
nfl_context_signals
nfl_edges
nfl_game_quality
nfl_prop_usage
nfl_war
nhl_eval
pickem
pickem_tiebreak
player_movement
preseason_signal
qb_handedness
qb_weather
split_miner
stable_filter_audit
stacking_eval
strategies
tennis_eval
weather_ftn_value
nhl_eval · 2
Do shift-derived ice time minutes improve shots-on-goal predictions? (2026-10-05)
Data: NHL shift charts 2021-22..2025-26 (6,936 games, ~5.3M shifts; 57 2024-25 games missing from the NHL feed) → player_games (TOI by EV/PP/PK + SOG). Regular season only. Scored 2022-23..2025-26: 169,298 player-games with ≥ 5 prior games. Every prediction is…
backtests/nhl_eval/FINDINGS_SOG_TOI.md
NHL Stacking v3 — De-Snooping Campaign Findings
Model: nhl_stacking_v3 (production's only currently-profitable model) Scope: honest, de-snooped performance measurement + hygiene-fix evaluation — NOT a search for new alpha Data: [redacted], 17,155 regular-season games, 2008–2024 (nhl_games join …
backtests/nhl_eval/FINDINGS.md
college_hockey_eval · 3
College hockey betting angles backlog (2026-09-29)
Context. The model loses to the market (FINDINGS_MODEL_VS_MARKET.md): log loss 0.625 vs 0.603. Fading it also loses. Anything that squeezes more prediction out of the same stats is dead. These angles are timing ones, or use information the line may not have.
backtests/college_hockey_eval/ANGLES_BACKLOG.md
Pre-registration: roster turnover vs the October market (2026-27)
Frozen 2026-09-29, before the season's first puck drop on 2026-10-02. The commit containing this file, forward_2026_team_ratings.json and roster_turnover_forward.py is the record. Nothing in them changes after 2026-10-02. Any later change goes in an "Amendment…
backtests/college_hockey_eval/PREREG_ROSTER_TURNOVER_2026.md
College hockey model vs the betting market (first comparison, 2026-09-29)
The market is better than the model, and the model adds nothing to it.
backtests/college_hockey_eval/FINDINGS_MODEL_VS_MARKET.md
qb_handedness · 1
fantasy_eval · 16
Do the Injury-Aware Multipliers Actually Improve Projections?
Validation of commit 02fba14, which replaced the projection engine's injury multiplier table using results from backtests/nfl_context_signals/FINDINGS.md.
backtests/fantasy_eval/FINDINGS_INJURY_MULTIPLIERS.md
If an Injured Player Suits Up, What Does He Score?
shipped
Follow-up to FINDINGS_INJURY_MULTIPLIERS.md, which shipped an expected-value injury multiplier and left the conditional one — "if he is active, what should I expect?" — as an aside: ~0.82 / 0.92 / 0.97 relative to healthy for DNP / Limited / Full. Nothing in t…
backtests/fantasy_eval/FINDINGS_INJURY_CONDITIONAL.md
Depth-chart multiplier damp — re-measured after the input fix (2026-09-12)
Verdict: keep 0.4. Mean error is flat across the damp; what moves is bias.
backtests/fantasy_eval/FINDINGS_DEPTH_DAMP.md
Snap Counts + Depth Charts Close ~30% of the ECR Gap
Date: 2026-07-19 · Season: 2025, weeks 7–17 · Harness: snap_depth_experiment.py → snap_depth_experiment_results.json · Seed: 20260719
backtests/fantasy_eval/FINDINGS_SNAP_DEPTH.md
Cross-player usage redistribution: the effect is real, the projection gain is not
rejected
Date: 2026-08-31 · Verdict: ❌ REJECTED — usage_redistribution = False
backtests/fantasy_eval/FINDINGS_USAGE_REDISTRIBUTION.md
Backup QB: shade the receivers, not the quarterback
shipped
Date: 2026-08-31 · Verdict: ✅ SHIPPED, receiving only
backtests/fantasy_eval/FINDINGS_BACKUP_QB.md
Train / dev / holdout: what the accuracy numbers are actually worth
Date: 2026-08-30 · Report: scripts/report_split_accuracy.py (re-runnable)
backtests/fantasy_eval/FINDINGS_EVAL_SPLIT.md
Two dead knobs: one was worth fixing, one was worth measuring and leaving off
shipped
Date: 2026-08-30 · Harness: evaluate_fantasy_mae.py --set <knob>=true
backtests/fantasy_eval/FINDINGS_PACE_AND_ROSTERS.md
Start/sit accuracy of the shipped surface
shipped
Date: 2026-08-30 · Harness: start_sit_live_accuracy.py (re-runnable)
backtests/fantasy_eval/FINDINGS_START_SIT_LIVE.md
Floor/ceiling were covering half the outcomes they claimed
Date: 2026-08-30 Verdict: ✅ FIXED — empirical residual quantiles replace the simulated band
backtests/fantasy_eval/FINDINGS_INTERVALS.md
The team+position "soup" is not a bug — it beats the clean matcher
rejected
Date: 2026-08-30 Harness: scripts/evaluate_fantasy_mae.py --set unified_player_matching=true|false Verdict: ❌ REJECTED — keep unified_player_matching = False
backtests/fantasy_eval/FINDINGS_PLAYER_MATCHING.md
Volume features: do they add anything on top of the snap re-ranker?
shipped
Date: 2026-08-30 Harness: backtests/fantasy_eval/volume_features_experiment.py (re-runnable) Verdict: ✅ SHIPPED — small but consistent incremental gain
backtests/fantasy_eval/FINDINGS_VOLUME_FEATURES.md
TD models: not unpredictable, just pointed at the wrong output
Date: 2026-08-30 Harness: backtests/fantasy_eval/td_probability_experiment.py (re-runnable) Verdict: ✅ PARTIAL — ships for rushing and receiving TDs, NOT for passing TDs
backtests/fantasy_eval/FINDINGS_TD_PROBABILITY.md
Our Fantasy Projections vs FantasyPros Expert Consensus (ECR)
Our side is the strict walk-forward regeneration from scripts/evaluate_fantasy_mae.py --dump-projections (2,438 player-weeks): the engine is fed one historical week at a time and may only use data from before the week it projects. Unlike the 2024 db rows, this…
backtests/fantasy_eval/FINDINGS_VS_ECR.md
The Close Band Can Be Improved (+2.5pp, ~half the ECR gap)
Date: 2026-07-20 · Harness: close_band_experiment.py · Sample: 12,510 close pairs of 56,046, 2025 weeks 5–17 · Seed: 20260719
backtests/fantasy_eval/FINDINGS_CLOSE_BAND.md
Start/Sit Decision Accuracy — 2025
Date: 2026-07-19 · Harness: start_sit_accuracy.py · Sample: 64,150 within-(week, position) pairs, 2025 weeks 5–17 · Seed: 20260719
backtests/fantasy_eval/FINDINGS_START_SIT.md
nfl_game_quality · 1
pickem · 6
How well does the public pick? — field baseline, 2014–2025
2026-09-13. field_baseline.py → field_baseline_results.json (gitignored). Source: NFL Pickwatch per-game pick shares, captured by scripts/nfl/capture_pickwatch.py into [redacted] (weekly cron, Tue 14:00 UTC). Lines and results come from nflverse, not…
backtests/pickem/FINDINGS_FIELD_BASELINE.md
Winning a 30-Person Confidence Pool
Date: 2026-07-20 · Harness: pool_strategy_sim.py → pool_strategy_sim_results.json · 6,000 simulated seasons per regime · Seed: 20260720
backtests/pickem/FINDINGS_POOL_2026.md
Testing "always pick the favorite" — and why a Monte Carlo can't do it
2026-08-12. calibration_test.py, 6,937 REG games, 1999–2025.
backtests/pickem/FINDINGS_CALIBRATION_2026.md
Season Prize vs Weekly Prize: Shoot for the Season
Date: 2026-07-20 · Harness: pool_ev_analysis.py → pool_ev_analysis_results.json · 4,000 simulated seasons per regime · Seed: 20260720
backtests/pickem/FINDINGS_EV_2026.md
Best Pick'em Strategy for 2026: Pick Every Favorite
Date: 2026-07-20 · Harness: strategy_search_2026.py → strategy_search_2026_results.json · Data: 2,119 REG games, 2018–2025 · Seed: 20260720
backtests/pickem/FINDINGS_STRATEGY_2026.md
Confidence Pick'em — Honest Rebuild (2018–2025)
Date: 2026-07-20 · Harness: backtest_pickem.py → backtest_pickem_results.json · Sample: 2,119 completed regular-season games · Seed: 20260720
backtests/pickem/FINDINGS.md
pickem_tiebreak · 1
nfl_war · 1
injury_impact · 1
player_movement · 2
What does transferring actually do? (college hockey)
2026-09-08. 5,494 skater and 307 goalie player-season pairs, 2020-2024, from college_hockey_player_game_stats joined to the transfer sheet ingested in e5746b5. A pair is a player who appears in season N and again in N+1, so both sides of a move are observed.
backtests/player_movement/FINDINGS_TRANSFERS.md
Detecting departures without an early-departure list
2026-09-08. We have no historical early-departure data, so the working premise was: read departures off game-1 rosters instead, accepting that we only learn them once the season starts and so cannot forecast with them.
backtests/player_movement/FINDINGS.md
elway_review · 1
kicker_dst · 1
fantasy_model_audit · 1
bet_timing · 1
middles · 1
leg_correlation · 1
stable_filter_audit · 1
ev_clv_prereg_2026_09 · 1
clv_line_adjustment · 1
ev_record_2026_08 · 1
weather_ftn_value · 1
qb_weather · 1
nfl_context_signals · 1
preseason_signal · 1
nfl_prop_usage · 2
NFL Hierarchical Usage Model — Evaluation vs Baselines
Date: 2026-07-14 · Seed: 20260714 · Harness: model_usage.py → results/usage_model_results.json + results/usage_model_predictions.parquet (88,724 walk-forward player-week-prop rows, 88 folds, 2021–2024 test seasons, 3,000 MC samples/row).
backtests/nfl_prop_usage/FINDINGS.md
NFL Prop Usage Model — Design Doc (Step 1: data + walk-forward harness)
Status: data layer + fold generator only. No model is fit yet. Deterministic seed 20260714. House rules apply: judge downstream models vs devigged closing lines, walk-forward only, cluster bootstraps by game_id / player_week_id, honest negative results are val…
backtests/nfl_prop_usage/DESIGN.md
hr_derby · 1
backtests · 3
College Eval Tracks 1+2 (NCAAB + CFB) — Synthesis, 2026-07-13
Verification method: independently re-opened the artifact files each build points to (walkforward_regen_results.json, replicate_spike_gate_result.json for CFB; production_clv_results.json for NCAAB) and diffed the build's prose numbers against the JSON on disk…
backtests/COLLEGE_EVAL_TRACKS_1_2_2026_07.md
Calibration Hygiene Campaign — Item 6 Synthesis
rejected
Date: 2026-07-12 | Scope: 5 tracks, adversarially checked | House framing: calibration is hygiene, not alpha — a change ships only if it clears a market-relative gate, never on ECE/log-loss improvement alone.
backtests/CALIBRATION_HYGIENE_2026_07.md
Upgrade Experiments 2026-07 — Item 5 Synthesis
Gate discipline reminder (house rule): walk-forward only, log-loss/MAE wins that worsen market-relative numbers don't ship, three clean FAILs would be a perfectly good outcome. All three experiments below cleared their pre-registered gates on the letter of the…
backtests/UPGRADE_EXPERIMENTS_2026_07.md
cfb_eval · 2
CFB-2b promotion spec — surface the walk-forward table on `/cfb/performance`
null
Status: SPEC ONLY — not applied. Ground rule for this task forbids editing production files or the route; this is the precise plan for the owner to apply.
backtests/cfb_eval/PROMOTION_SPEC_2b.md
CFB-2a promotion spec — market arm in `scripts/calibrate_models.py`
Status: SPEC ONLY — not applied. Ground rule for this task forbids editing production files; this is the precise diff for the owner to apply by hand.
backtests/cfb_eval/PROMOTION_SPEC.md
ncaab_eval · 1
tennis_eval · 1
split_miner · 1
strategies · 1
stacking_eval · 4
Tennis consensus-stacking evaluation — REJECT (2026-07-04)
rejected
A full historical backtest was possible: [redacted]::tennis_matches (81,622 matches 2010-2026 with embedded closing odds: market average, Pinnacle, best price) + tennis_match_elo (point-in-time pre-match Elo, through 2024-11-17). Joined usable set: 69,090 …
backtests/stacking_eval/tennis/REPORT.md
NCAAB consensus-stacking evaluation — REJECT (2026-07-04)
rejected
The best-prior candidate (soft small-conference markets), tested on our largest capture. Verdict: consensus wins everything.
backtests/stacking_eval/ncaab/REPORT.md
MLB consensus-stacking pilot — INSUFFICIENT DATA, leaning REJECT (2026-07-04)
rejected
Question: does blending the pitcher-adjusted Elo v2 prob into the devigged closing consensus (NHL v3 recipe) beat consensus-only?
backtests/stacking_eval/mlb/REPORT.md
NFL consensus-stacking evaluation — REJECT (2026-07-04)
rejected
Question: does the NHL v3 "top-down" recipe (tiny logistic on devigged closing consensus + small orthogonal features) beat the NFL closing moneyline?
backtests/stacking_eval/nfl/REPORT.md
kelly_sim · 1
nfl_edges · 1
docs · 9
Strategy 002: 12-1 Cross-Sectional Momentum
Jegadeesh & Titman, Returns to Buying Winners and Selling Losers (Journal of Finance, 1993). Still one of the most-replicated equity anomalies.
backtests/docs/002_momentum_12_1.md
Strategy 009: Insider Cluster Buying
Cohen, Malloy & Pomorski, Decoding Inside Information (Journal of Finance, 2012). Also Lakonishok & Lee (2001), Jeng et al. (2003).
backtests/docs/009_insider_cluster.md
Strategy 008: Sell-in-May Calendar Effect
Folk wisdom + Bouman & Jacobsen, The Halloween Indicator, "Sell-in-May-and-Go-Away": Another Puzzle (American Economic Review, 2002). Subsequent decades of replication.
backtests/docs/008_sell_in_may.md
Strategy 007: Short-Term Reversal (1-week loser-winner)
Jegadeesh (1990), Lehmann (1990) — the original short-horizon reversal papers. The mirror image of 12-1 momentum: short-term returns reverse at 1-week/1-month horizons, while medium-term (12-1) returns continue.
backtests/docs/007_short_term_reversal.md
Strategy 006: Net-Net (Graham Deep Value)
Benjamin Graham, The Intelligent Investor (1949) and Security Analysis (1934). The original quantitative value strategy. Tweedy Browne, Walter Schloss, and others used variants for decades.
backtests/docs/006_net_net.md
Strategy 005: VIX Term Structure
Multiple — vol-arbitrage literature 2010s. Notable: Simon (2014), Donninger (2015), the "Hayek" / "Boggs" VIX-term-structure papers. Popularized in retail circles via SVXY / VXX strategies after the 2011 launch of VIX ETPs.
backtests/docs/005_vix_term_structure.md
Strategy 004: Piotroski F-Score
Joseph Piotroski, Value Investing: The Use of Historical Financial Statement Information to Separate Winners from Losers (Journal of Accounting Research, 2000).
backtests/docs/004_piotroski.md
Strategy 003: Post-Earnings Announcement Drift (PEAD)
Bernard & Thomas, Post-Earnings-Announcement Drift: Delayed Price Response or Risk Premium? (1989). Replicated and refined for ~40 years.
backtests/docs/003_pead.md
Strategy 001: Magic Formula (Greenblatt)
Joel Greenblatt, The Little Book That Beats the Market (2006).
backtests/docs/001_magic_formula.md