New to these reports? Start here
  • Dotted-underlined words have a plain-English definition — hover or tap them. Every term is also on the glossary page.
  • "Null" means we found nothing, not that something broke. Most reports here are negative results, on purpose — knowing an idea doesn't work is the point.
  • Two questions get asked separately. First, is the effect real? Second, is it already priced into the betting odds? An effect can be completely real and still useless to bet on.
  • A "calibration" row is a self-check. It runs the same method on something already known to be true. If that fails, the whole report is unreliable — so it's reported alongside the findings.
  • If a confidence interval includes zero, the real effect might be nothing at all, so no claim gets made.

Floor/ceiling were covering half the outcomes they claimed

Date: 2026-08-30 Verdict: ✅ FIXED — empirical residual quantiles replace the simulated band

The problem

floor and ceiling are the 10th and 90th percentiles, so 80% of actual scores should land between them. On all 16,013 graded player-weeks they contained 50.2% — and lopsidedly:

measured target
below floor 32.7% 10%
inside 50.2% 80%
above ceiling 17.0% 10%

Why — three things at once, none of them simulation noise

The band came from resampling a symmetric normal around the projected stat line. That gets three separate things wrong:

  1. Width. Simply too narrow — it needed roughly 70% more.
  2. Scale. The real error spread grows with the projection. The 10th-to-90th residual range is ~8 points wide for a 2-point projection and ~25 points for a 20-point one. One width cannot serve both.
  3. Shape. Scoring is right-skewed and floored at zero, and the asymmetry flips with projection level:
projection residual p10 residual p90
0–3 −2.3 +6.1
9–12 −8.9 +9.6
20+ −16.2 +8.8

A 2-point player can boom; a 20-point player can be hurt or benched but cannot vastly exceed his role. A symmetric band cannot express either end.

⚠️ More Monte Carlo samples fixes none of this. Sampling noise was ±0.68 points per endpoint; the calibrationcalibrationA deliberate sanity check on the method itself: run it on something already known to be true. If it fails to detect the known thing, the method is broken and its other results mean nothing. error was ~3.5. Adding samples makes a wrong number more precise. (It also never came from simulation_count — it was a hardcoded 100-draw loop resampling an already-simulated mean/std.)

The fix

Take the empirical 10th/90th percentiles of actual − projection, split by position and projection level, fit on prior seasons only, and interpolate between levels. Refit each season on everything before it:

season old below/in/above calibrated
2022 33.6 / 50.0 / 16.5 11.3 / 77.6 / 11.1
2023 31.3 / 51.0 / 17.7 8.8 / 81.2 / 10.1
2024 30.3 / 50.4 / 19.3 9.9 / 78.6 / 11.5
2025 34.4 / 49.6 / 16.0 12.3 / 78.2 / 9.5

Target is 10 / 80 / 10. Coverage lands in 77.6–81.2% every year, and both tails sit near 10% instead of 33/17.

Shipped

[redacted], on by default (calibrated_intervals), applied after the top-end shrink so it anchors to the final PPRpoints per receptionA fantasy scoring format that awards a point for every catch, which raises the value of high-volume receivers. value. Returns None without enough history and the simulated band stands — a miscalibrated band is bad, a fabricated one is worse. LEGACY_ENGINE_CONFIG sets it False so the evaluation harness can still reproduce the old numbers.

Point projections are unchanged — verified 0 differences on every deterministic field against a pre-change regeneration. Only the band moved (247 of 256 players).

Tests: scripts/unit/test_interval_calibration.py (9), including an out-of-sampleout-of-sampleTested on data that was not used to build or tune the idea. This is the honest test; results on the data you built with are almost always flattering. coverage assertion so a future change cannot silently de-calibrate this again.