Research Library
Every findings document in the repo, rendered. Sourced directly from the
markdown under backtests/ and docs/ — nothing to
register, a new file shows up here automatically. Negative results are kept
deliberately: most of what's here is a hypothesis that didn't survive.
38 documents are private and not shown.
All areas
backtests
cfb_eval
docs
fantasy_eval
hr_derby
kelly_sim
ncaab_eval
nfl_context_signals
nfl_edges
nfl_prop_usage
nhl_eval
pickem
preseason_signal
qb_weather
split_miner
stacking_eval
strategies
tennis_eval
weather_ftn_value
weather_ftn_value · 1
qb_weather · 1
fantasy_eval · 5
Do the Injury-Aware Multipliers Actually Improve Projections?
shipped
Validation of commit 02fba14, which replaced the projection engine's injury multiplier table using results from backtests/nfl_context_signals/FINDINGS.md.
backtests/fantasy_eval/FINDINGS_INJURY_MULTIPLIERS.md
Our Fantasy Projections vs FantasyPros Expert Consensus (ECR)
Our side is the strict walk-forward regeneration from scripts/evaluate_fantasy_mae.py --dump-projections (2,438 player-weeks): the engine is fed one historical week at a time and may only use data from before the week it projects. Unlike the 2024 db rows, this…
backtests/fantasy_eval/FINDINGS_VS_ECR.md
The Close Band Can Be Improved (+2.5pp, ~half the ECR gap)
Date: 2026-07-20 · Harness: close_band_experiment.py · Sample: 12,510 close pairs of 56,046, 2025 weeks 5–17 · Seed: 20260719
backtests/fantasy_eval/FINDINGS_CLOSE_BAND.md
Start/Sit Decision Accuracy — 2025
Date: 2026-07-19 · Harness: start_sit_accuracy.py · Sample: 64,150 within-(week, position) pairs, 2025 weeks 5–17 · Seed: 20260719
backtests/fantasy_eval/FINDINGS_START_SIT.md
Snap Counts + Depth Charts Close ~30% of the ECR Gap
Date: 2026-07-19 · Season: 2025, weeks 7–17 · Harness: snap_depth_experiment.py → snap_depth_experiment_results.json · Seed: 20260719
backtests/fantasy_eval/FINDINGS_SNAP_DEPTH.md
nfl_context_signals · 1
preseason_signal · 1
pickem · 5
Winning a 30-Person Confidence Pool
Date: 2026-07-20 · Harness: pool_strategy_sim.py → pool_strategy_sim_results.json · 6,000 simulated seasons per regime · Seed: 20260720
backtests/pickem/FINDINGS_POOL_2026.md
Testing "always pick the favorite" — and why a Monte Carlo can't do it
2026-08-12. calibration_test.py, 6,937 REG games, 1999–2025.
backtests/pickem/FINDINGS_CALIBRATION_2026.md
Season Prize vs Weekly Prize: Shoot for the Season
Date: 2026-07-20 · Harness: pool_ev_analysis.py → pool_ev_analysis_results.json · 4,000 simulated seasons per regime · Seed: 20260720
backtests/pickem/FINDINGS_EV_2026.md
Best Pick'em Strategy for 2026: Pick Every Favorite
Date: 2026-07-20 · Harness: strategy_search_2026.py → strategy_search_2026_results.json · Data: 2,119 REG games, 2018–2025 · Seed: 20260720
backtests/pickem/FINDINGS_STRATEGY_2026.md
Confidence Pick'em — Honest Rebuild (2018–2025)
Date: 2026-07-20 · Harness: backtest_pickem.py → backtest_pickem_results.json · Sample: 2,119 completed regular-season games · Seed: 20260720
backtests/pickem/FINDINGS.md
nfl_prop_usage · 2
NFL Hierarchical Usage Model — Evaluation vs Baselines
Date: 2026-07-14 · Seed: 20260714 · Harness: model_usage.py → results/usage_model_results.json + results/usage_model_predictions.parquet (88,724 walk-forward player-week-prop rows, 88 folds, 2021–2024 test seasons, 3,000 MC samples/row).
backtests/nfl_prop_usage/FINDINGS.md
NFL Prop Usage Model — Design Doc (Step 1: data + walk-forward harness)
Status: data layer + fold generator only. No model is fit yet. Deterministic seed 20260714. House rules apply: judge downstream models vs devigged closing lines, walk-forward only, cluster bootstraps by game_id / player_week_id, honest negative results are val…
backtests/nfl_prop_usage/DESIGN.md
hr_derby · 1
backtests · 3
College Eval Tracks 1+2 (NCAAB + CFB) — Synthesis, 2026-07-13
Verification method: independently re-opened the artifact files each build points to (walkforward_regen_results.json, replicate_spike_gate_result.json for CFB; production_clv_results.json for NCAAB) and diffed the build's prose numbers against the JSON on disk…
backtests/COLLEGE_EVAL_TRACKS_1_2_2026_07.md
Calibration Hygiene Campaign — Item 6 Synthesis
rejected
Date: 2026-07-12 | Scope: 5 tracks, adversarially checked | House framing: calibration is hygiene, not alpha — a change ships only if it clears a market-relative gate, never on ECE/log-loss improvement alone.
backtests/CALIBRATION_HYGIENE_2026_07.md
Upgrade Experiments 2026-07 — Item 5 Synthesis
Gate discipline reminder (house rule): walk-forward only, log-loss/MAE wins that worsen market-relative numbers don't ship, three clean FAILs would be a perfectly good outcome. All three experiments below cleared their pre-registered gates on the letter of the…
backtests/UPGRADE_EXPERIMENTS_2026_07.md
cfb_eval · 2
CFB-2b promotion spec — surface the walk-forward table on `/cfb/performance`
null
Status: SPEC ONLY — not applied. Ground rule for this task forbids editing production files or the route; this is the precise plan for the owner to apply.
backtests/cfb_eval/PROMOTION_SPEC_2b.md
CFB-2a promotion spec — market arm in `scripts/calibrate_models.py`
Status: SPEC ONLY — not applied. Ground rule for this task forbids editing production files; this is the precise diff for the owner to apply by hand.
backtests/cfb_eval/PROMOTION_SPEC.md
ncaab_eval · 1
tennis_eval · 1
nhl_eval · 1
split_miner · 1
strategies · 1
stacking_eval · 4
Tennis consensus-stacking evaluation — REJECT (2026-07-04)
rejected
A full historical backtest was possible: [redacted]::tennis_matches (81,622 matches 2010-2026 with embedded closing odds: market average, Pinnacle, best price) + tennis_match_elo (point-in-time pre-match Elo, through 2024-11-17). Joined usable set: 69,090 …
backtests/stacking_eval/tennis/REPORT.md
NCAAB consensus-stacking evaluation — REJECT (2026-07-04)
rejected
The best-prior candidate (soft small-conference markets), tested on our largest capture. Verdict: consensus wins everything.
backtests/stacking_eval/ncaab/REPORT.md
MLB consensus-stacking pilot — INSUFFICIENT DATA, leaning REJECT (2026-07-04)
rejected
Question: does blending the pitcher-adjusted Elo v2 prob into the devigged closing consensus (NHL v3 recipe) beat consensus-only?
backtests/stacking_eval/mlb/REPORT.md
NFL consensus-stacking evaluation — REJECT (2026-07-04)
rejected
Question: does the NHL v3 "top-down" recipe (tiny logistic on devigged closing consensus + small orthogonal features) beat the NFL closing moneyline?
backtests/stacking_eval/nfl/REPORT.md
kelly_sim · 1
nfl_edges · 1
docs · 9
Strategy 002: 12-1 Cross-Sectional Momentum
Jegadeesh & Titman, Returns to Buying Winners and Selling Losers (Journal of Finance, 1993). Still one of the most-replicated equity anomalies.
backtests/docs/002_momentum_12_1.md
Strategy 009: Insider Cluster Buying
Cohen, Malloy & Pomorski, Decoding Inside Information (Journal of Finance, 2012). Also Lakonishok & Lee (2001), Jeng et al. (2003).
backtests/docs/009_insider_cluster.md
Strategy 008: Sell-in-May Calendar Effect
Folk wisdom + Bouman & Jacobsen, The Halloween Indicator, "Sell-in-May-and-Go-Away": Another Puzzle (American Economic Review, 2002). Subsequent decades of replication.
backtests/docs/008_sell_in_may.md
Strategy 007: Short-Term Reversal (1-week loser-winner)
Jegadeesh (1990), Lehmann (1990) — the original short-horizon reversal papers. The mirror image of 12-1 momentum: short-term returns reverse at 1-week/1-month horizons, while medium-term (12-1) returns continue.
backtests/docs/007_short_term_reversal.md
Strategy 006: Net-Net (Graham Deep Value)
Benjamin Graham, The Intelligent Investor (1949) and Security Analysis (1934). The original quantitative value strategy. Tweedy Browne, Walter Schloss, and others used variants for decades.
backtests/docs/006_net_net.md
Strategy 005: VIX Term Structure
Multiple — vol-arbitrage literature 2010s. Notable: Simon (2014), Donninger (2015), the "Hayek" / "Boggs" VIX-term-structure papers. Popularized in retail circles via SVXY / VXX strategies after the 2011 launch of VIX ETPs.
backtests/docs/005_vix_term_structure.md
Strategy 004: Piotroski F-Score
Joseph Piotroski, Value Investing: The Use of Historical Financial Statement Information to Separate Winners from Losers (Journal of Accounting Research, 2000).
backtests/docs/004_piotroski.md
Strategy 003: Post-Earnings Announcement Drift (PEAD)
Bernard & Thomas, Post-Earnings-Announcement Drift: Delayed Price Response or Risk Premium? (1989). Replicated and refined for ~40 years.
backtests/docs/003_pead.md
Strategy 001: Magic Formula (Greenblatt)
Joel Greenblatt, The Little Book That Beats the Market (2006).
backtests/docs/001_magic_formula.md