New to these reports? Start here
- Dotted-underlined words have a plain-English definition — hover or tap them. Every term is also on the glossary page.
- "Null" means we found nothing, not that something broke. Most reports here are negative results, on purpose — knowing an idea doesn't work is the point.
- Two questions get asked separately. First, is the effect real? Second, is it already priced into the betting odds? An effect can be completely real and still useless to bet on.
- A "calibration" row is a self-check. It runs the same method on something already known to be true. If that fails, the whole report is unreliable — so it's reported alongside the findings.
- If a confidence interval includes zero, the real effect might be nothing at all, so no claim gets made.
The CLV "line-match bug" is a missing-data problem, not a maths error
2026-08-24. Investigated because mlb/totals and nba/spreads showed
absurd CLVclosing line valueWhether you got a better price than the market settled at. Widely used as a faster signal of skill than profit, which takes ages to measure. capture coverage (39% and 18%) together with anomalously small
dispersion — nba/spreads sd 0.93 against nba/h2h sd 3.31 in the same sport.
Low coverage plus compressed variance is the signature of a selection effect,
not a quiet market.
What clv_unreliable actually is
It is exactly a line-movement flag. Over Feb–Aug 2026, spread/total alerts:
| n | mean CLV | sd | line matched the close exactly | |
|---|---|---|---|---|
clv_unreliable = 0 |
796 | −0.20% | 1.84 | 100.0% |
clv_unreliable = 1 |
1,821 | −0.35% | 3.99 | 0.0% (mean |diff| 1.12 pts) |
| all | 2,617 | −0.30% | 3.48 | — |
The partition is perfect: reliable ⟺ the line did not move.
So filtering on clv_unreliable = 0 — which every analysis in this project has
done, including the segment table in the CLV pre-registration — does not remove
bad measurements. It removes every game whose line moved, which is 70% of
spread/total alerts. Lines move when information arrives, so this conditions
the sample on the uninformative subset. The consequences are exactly what was
observed: variance more than halves (3.99 → 1.84) and the mean is pulled toward
zero (−0.35% → −0.20%).
This is missing-not-at-random. No sample size repairs it.
The flag itself is correct and should not be removed. An Over 8.5 genuinely is a different bet from an Over 9.5, and grading one against the other would be wrong — which is what the flag was added to prevent. The error is downstream, in treating "unreliable" as "discard" rather than "needs pricing".
How much is recoverable, and how
At the last snapshot before kickoff, for the 1,821 flagged alerts:
| n | share | |
|---|---|---|
| some book still quotes the alert's exact line | 816 | 44.8% |
| …of which Pinnacle | 48 | 2.6% |
| line is gone everywhere — needs point pricing | 1,005 | 55.2% |
So roughly 45% is recoverable with no modelling at all: widen the closing lookup beyond Pinnacle to any book still quoting that exact line at close, and devigdevigRemoving the bookmaker's cut from odds to recover the market's actual implied probability. that book's two sides. The books that supply it are mostly soft (FanDuel 595, DraftKings 568, BetMGM 492, BetRivers 425) — a soft book's close is a worse estimate of true probability than Pinnacle's, so these must be recorded as a separate quality tier, not silently pooled.
⚠ An earlier version of this analysis reported "100% recoverable" by asking whether any book quoted the line at any pre-kickoff snapshot. That is wrong: an early price is not a closing price. Restricting to the closing snapshot gives the 44.8% above.
The remaining 55% requires pricing the point difference.
Pricing a point: cross-book identification works only where lines are pinned
At a single instant, different books post different lines on the same game — the same event at different points with no information difference. Fitting a through-origin slope of devigged probability on point difference, pooled within (game, snapshot) groups so each game's baseline differences out:
| segment | distinct lines posted | fitted half-point | verdict |
|---|---|---|---|
| nhl/spreads | 2 | 7.02pp | ✅ clean |
| nhl/totals | 3 | 6.11pp | ✅ clean |
| mlb/spreads | 10 | 5.01pp | ✅ clean |
| mlb/totals | 12 | 4.03pp | ✅ clean |
| ncaab/spreads | — | 1.00pp | ⚠ attenuated |
| nfl/spreads | 55 | 0.60pp | ❌ ~3x too small |
| nba/spreads | 32 | 0.51pp | ❌ ~2x too small |
| nba/totals | 35 | 0.15pp | ❌ ~7x too small |
| nfl/totals | 34 | 0.14pp | ❌ ~10x too small |
The fitted slope tracks inversely with how freely books move the line, and the
mechanism is identifiable. If book A thinks the total is 8 and prices Over 8
near fair 50%, while book B thinks it is 9 and prices Over 9 near fair 50%, then
dPoint = −1 and dProb = 0: the slope is driven to zero. Books that move the
line re-centre the price to their own opinion, so the point difference
carries no probability difference.
Where books cannot move the line — the MLB run line and NHL puck line, both pinned at ±1.5 — disagreement must be expressed in the price instead, the contrast survives, and the fit lands on conventional values (half-run 4.0pp, half-goal 6.1pp). That agreement with independent market convention is the validation.
So cross-book identification is usable for MLB and NHL only. NFL, NBA and NCAAB need the empirical distribution of the final margin relative to the closing lineclosing lineThe final odds right before a game starts. It reflects everything the betting market knows, which makes it the hardest benchmark to beat., which also captures NFL's key numbers at 3 and 7 that no single slope can represent.
Data problem found on the way
⚠ [redacted] :: nba_sbr_odds.close_home_spread is corrupt.
35.3% of its values are NBA totals sitting in the spread column (215.0, 217.5,
218.5 …), and every value is unsigned — the p0 is +1.0, so no home underdog
exists anywhere in 12,036 rows. Unusable as a closing-spread source. Not
investigated further here; flagged because anything reading that column is
getting nonsense.
Result of shipping the exact-line lookup
_lookup_exact_line_close was implemented and replayed over the 1,821 discarded
alerts. It recovered 756 (41.5%), against the 44.8% availability estimate —
the gap is books quoting only one side, which cannot be devigged.
First: is a soft-book close comparable to Pinnacle's? 98% of recoveries come from soft books, so a naive comparison could just be measuring book quality. On 761 games where Pinnacle and a soft book both quoted the same line, the devigged difference is −0.025pp [−0.091, +0.040] — indistinguishable from zero. The recovered numbers are therefore comparable, and no tier correction is needed. (The tier is still recorded, so this can be re-checked.)
The recovered rows are five times worse than the ones we kept:
| n | mean CLV | 95% CI95% confidence intervalThe range the true value is plausibly in. If this range includes zero, we cannot rule out that the real effect is nothing at all. | |
|---|---|---|---|
| line did not move — what we reported | 796 | −0.20% | [−0.33, −0.07] |
| line moved, close recovered — was dropped | 756 | −1.04% | [−1.15, −0.92] |
| combined — the honest figure | 1,552 | −0.61% | [−0.70, −0.52] |
So excluding moved lines made spread/total CLV look 3x better than it is. The interpretation is direct: when the line moves off your number you are systematically on the wrong side of it — being picked off, not capturing value.
⚠ −0.61% is an optimistic bound. Recovery is itself selected toward small moves: recovered alerts average 0.67 points of movement, the 1,065 still unmeasured average 1.45. CLV worsens monotonically with move size (−0.20% at zero, −1.04% at 0.67), so the unmeasured tail is likely worse still. Closing that gap needs the point-pricing work.
Tier 3: pricing the move when no book holds the line
Where cross-book fitting fails, the ground truth is the empirical distribution of
the result against the closing line (fit_residuals.py). NBA closing spreads
were recovered from the corrupt nba_sbr_odds column by keeping |value| ≤ 30 and
restoring the sign from the moneyline — validated at a 48.65% home cover rate,
residual mean −0.40, sd 12.52 against an 11–12 convention. NFL came from
nflreadpy (3,028 games, 100% line coverage).
The empirical half-point values confirm the attenuation diagnosis quantitatively:
| segment | empirical | cross-book | ratio |
|---|---|---|---|
| nba/spreads | 1.84pp | 0.51pp | 3.6× |
| nfl/spreads | 2.39pp | 0.60pp | 4.0× |
| nfl/totals | 1.72pp | 0.14pp | 12.1× |
Validated against known truth. For 338 alerts whose true exact-line close was independently recoverable, predicting it from the moved-line close:
| n | biasbiasWhether the misses lean consistently one way. A projection can have a good average error size but still be biased if it is almost always too high. Bias is often the more fixable problem. | MAEmean absolute errorAverage size of the miss, ignoring direction. If a projection is off by 3 one week and -5 the next, the MAE is 4. Lower is better. | |
|---|---|---|---|
| no adjustment (previous behaviour) | 338 | +0.217pp | 4.471pp |
| with point adjustment | 338 | −0.137pp | 1.217pp |
A 72.8% error reduction, on tables fitted from 2015–2025 data against 2026 alerts. Gains concentrate where a point is expensive — nhl/totals 6.26→1.24pp, mlb/totals 4.05→1.10pp — while nba/spreads improves least (1.39→1.25) because an NBA half-point is only worth ~1.8pp to begin with.
The honest number, all three tiers
| tier | n | mean CLV | 95% CI |
|---|---|---|---|
| exact — line never moved | 796 | −0.20% | [−0.33, −0.07] |
| exact — recovered from another book | 756 | −1.04% | [−1.15, −0.92] |
| point-adjusted | 546 | −2.43% | [−2.95, −1.90] |
| all measured | 2,098 | −1.08% | [−1.24, −0.93] |
Coverage 30.4% → 80.2%. The reported −0.20% was the best 30% of a distribution whose true mean is −1.08% — five times better than reality.
The pattern is monotonemonotoneConsistently moving one direction as the input increases. A real dose-response effect should be monotone; if more of the cause does not mean more of the effect, the pattern is suspect. in line movement: no move −0.20%, small move −1.04%, large move −2.43%. That is the signature of being picked off. When the market moves off your number, you are on the wrong side of it, and the further it moves the worse you did.
Blocked segments now clear their unblock criterion
Pre-specified as coverage ≥85% and dispersion within 2× the same sport's h2h:
| segment | coverage | sd | ratio vs h2h | unblock |
|---|---|---|---|---|
| nba/spreads | 100% | 5.94 | 1.58 | YES |
| mlb/totals | 100% | 2.73 | 1.14 | YES |
nba/spreads went from sd 0.93 at 18% coverage to sd 5.94 at 100% — the
compressed-variance pathology that prompted the whole investigation is gone.
NCAAB point markets remain unpriceable (no residual table, not pinned) at ~65%
coverage, but ncaab/spreads and ncaab/totals are both disabled anyway.
Status
- [x] Diagnosis confirmed
- [x] Recoverable fraction quantified (44.8% available, 41.5% realised)
- [x] Point values fitted and validated for MLB/NHL, rejected for NFL/NBA/NCAAB
- [x] Exact-line lookup shipped, tier-tagged, soft-vs-sharp bias measured at ~0
- [x] Empirical residual tables for NBA/NFL; adjustment validated at −72.8% MAE
- [x] Honest spread/total CLV re-measured: −0.20% → −1.08%, coverage 30% → 80%
- [x]
mlb/totalsandnba/spreadsunblocked in the pre-registration - [ ] NCAAB point markets remain unpriceable (both disabled, so not blocking)