New to these reports? Start here
- Dotted-underlined words have a plain-English definition — hover or tap them. Every term is also on the glossary page.
- "Null" means we found nothing, not that something broke. Most reports here are negative results, on purpose — knowing an idea doesn't work is the point.
- Two questions get asked separately. First, is the effect real? Second, is it already priced into the betting odds? An effect can be completely real and still useless to bet on.
- A "calibration" row is a self-check. It runs the same method on something already known to be true. If that fails, the whole report is unreliable — so it's reported alongside the findings.
- If a confidence interval includes zero, the real effect might be nothing at all, so no claim gets made.
Start/sit accuracy of the shipped surface
Date: 2026-08-30 · Harness: start_sit_live_accuracy.py (re-runnable)
Scores what /fantasy/start-sit actually does today — the projection ordering
plus the deployed snap/depth/volume re-ranker — on the rebuilt walk-forwardwalk-forwardEvaluating week by week using only what was knowable before each week, mimicking how the model would actually have been used at the time.
history for 2025. 70,504 within-(week, position) pairs. A call is correct when
the higher-ranked player actually outscores the other; actual ties dropped.
| projected gap | pairs | projection | + re-ranker | expert consensus (ECRexpert consensus rankingFantasyPros' averaged ranking across many fantasy analysts. A strong benchmark, and hard to beat.) |
|---|---|---|---|---|
| coin-flip (<2 apart) | 16,263 | 54.9% | 57.4% | 60.3% |
| close (2–5) | 20,904 | 64.1% | 64.9% | 66.5% |
| clear (5–10) | 23,252 | 76.0% | 76.0% | 76.7% |
| blowout (10+) | 10,085 | 87.7% | 87.7% | 88.0% |
| all | 70,504 | 69.3% | 70.1% | 71.5% |
Re-ranker gain on coin-flip calls: +2.49%, 95% CI95% confidence intervalThe range the true value is plausibly in. If this range includes zero, we cannot rule out that the real effect is nothing at all. [+0.92%, +4.30%] (bootstrapbootstrapRe-running a calculation on thousands of resampled versions of the data to see how much the answer wobbles. The spread of those answers becomes the confidence interval. clustered by week). That independently reproduces the +2.51pp the original snap/depth study predicted, on different data and through the shipped code path rather than the experiment's.
Read it honestly. The re-ranker only moves the decisions it was supposed to move — coin-flips and, slightly, the close band — and does nothing at 5+ point gaps, which is correct: those are not decisions anyone agonises over. Aggregate accuracy (70.1%) is flattering and mostly measures blowouts.
ECR still wins at every gap. On the calls that matter it is 60.3% against our 57.4%. The re-ranker closed roughly 46% of that gap (5.4pp → 2.9pp) but did not erase it. Consistent with FINDINGS_VS_ECR.md: expert consensus out-ranks this engine, and the start/sit surface should keep leading with ECR.