New to these reports? Start here
- Dotted-underlined words have a plain-English definition — hover or tap them. Every term is also on the glossary page.
- "Null" means we found nothing, not that something broke. Most reports here are negative results, on purpose — knowing an idea doesn't work is the point.
- Two questions get asked separately. First, is the effect real? Second, is it already priced into the betting odds? An effect can be completely real and still useless to bet on.
- A "calibration" row is a self-check. It runs the same method on something already known to be true. If that fails, the whole report is unreliable — so it's reported alongside the findings.
- If a confidence interval includes zero, the real effect might be nothing at all, so no claim gets made.
EV Alert Pipeline — Record Review, August 2026
Review of the live EV alert record and the config changes made as a result.
Data: [redacted], 12,821 alerts (Feb–Aug 2026), 12,718 settled
win/loss. This is the deduped store (write-time dedup added 2026-07;
history collapsed 34.7k → 12.6k), so rows are one-per-pick, not
duplicate-weighted.
ROIreturn on investmentProfit as a percentage of the money wagered. +2% means $2 profit per $100 bet. is flat 1u at the actual alert odds, pushes excluded.
Headline
| Overall ROI | −1.94% [−3.62, −0.26] |
| Settled alerts | 12,718 |
| Net | −246.5 units |
A small but statistically detectable bleed — the CI excludes zero.
⚠ Recent months are uninformative — but the reason is SEASONALITY, not the budget rework. An earlier draft of this review attributed the volume drop to the August rework. That was wrong. Volume by month and sport:
| Month | MLB | NBA | NCAAB | NHL | Total |
|---|---|---|---|---|---|
| 2026-02 | – | 2,042 | 839 | 7 | 2,888 |
| 2026-03 | – | 4,395 | 559 | 538 | 5,492 |
| 2026-04 | – | 594 | – | 103 | 697 |
| 2026-05 | 52 | 2,152 | – | 601 | 2,805 |
| 2026-06 | 329 | 168 | – | 60 | 722 |
| 2026-07 | 98 | – | – | – | 98 |
| 2026-08 | 107 | – | – | – | 119 |
July–August is MLB-only because MLB is the only sport playing. NBA, NCAAB and NHL all ended. The pipeline is in its seasonal trough, not a degraded state. August's −15% on n=109 (CI spanning 36 points) still means nothing, but "volume collapsed" was the wrong diagnosis.
Where the loss actually is
Four markets accounted for 234 of the 246 units lost, and all four were still enabled at the time of this review:
| Market | n | ROI | 95% CI95% confidence intervalThe range the true value is plausibly in. If this range includes zero, we cannot rule out that the real effect is nothing at all. | Units |
|---|---|---|---|---|
| NBA totals | 525 | −22.46% | [−30.5, −14.4] | −117.9 |
| NHL totals | 163 | −31.83% | [−45.4, −18.3] | −51.9 |
| NCAAB spreads | 592 | −8.67% | [−16.7, −0.7] | −51.3 |
| NCAAB totals | 48 | −26.80% | [−53.5, −0.1] | −12.9 |
Counterfactual with those four removed: −1.94% → −0.11%, net −246.5u → −12.5u. Everything else is roughly break-even.
Why this is not just post-hoc subgroup mining
Selecting losers by inspecting their record is exactly the grid-argmax trap this project has documented as skill-less, so it deserves scepticism. Two things override it here:
- The June 2026 ROI audit already flagged "NBA totals as the big leaks." This is a replication of a previously-identified finding that was never actioned — not a fresh discovery.
- NBA totals at −22.5% with an upper CI bound of −14.4% over 525 bets is far outside what chance produces across ~11 markets tested.
Confidence is not uniform: NBA/NHL totals are strong; NCAAB spreads is moderate; NCAAB totals (n=48, CI upper bound −0.07%) is marginal and was disabled mainly because its volume is tiny and the cost of being wrong is near zero.
The EV number itself is not working
| EV bucket | n | ROI | 95% CI |
|---|---|---|---|
| 0–2% | 5,559 | −1.56% | [−3.99, +0.87] |
| 2–4% | 4,178 | −1.95% | [−4.91, +1.02] |
| 4–6% | 1,305 | +1.24% | [−4.18, +6.66] |
| 6–10% | 564 | +0.14% | [−8.29, +8.57] |
| 10–20% | 371 | +2.10% | [−8.26, +12.46] |
| 20%+ | 696 | −15.67% | [−23.57, −7.76] |
corr(predicted EV, realised P&L) = −0.043 across 12,718 alerts. Predicted EV
carries essentially no information about realised profit.
⚠ Correction: the EV cap already exists and already works
The initial recommendation from this review was "deploy an EV cap." That was
wrong — EVFinder.MAX_EV_PCT = 8.0 is in the code and is applied. Checking
the data confirms it works: the last market_consensus alert above 8% EV was in
February 2026.
Decomposing the 20%+ bucket by source shows the penalty is not a general law but two specific model sources:
| Source at EV ≥ 20% | n | ROI | 95% CI |
|---|---|---|---|
nba_model_v2 |
530 | −17.96% | [−26.41, −9.50] |
nhl_model_v1 |
71 | −28.17% | [−47.74, −8.60] |
tennis_elo_v1 |
94 | +4.24% | [−25.56, +34.05] |
Both offenders are already handled: nba_model_v2 is capped at 4.0 EV
(alert_tracker.py), and nhl_model_v1 is retired. No further EV cap was
added.
Specifically, tennis_elo_v1 was left uncapped despite emitting 28–35% EV
alerts, because its high-EV subset is the one that is not losing (+4.24%) and
tennis overall is the only positive segment (+14.87%, n=113). Capping it would
have been acting against the evidence in order to satisfy a rule derived from a
different, already-fixed source. Tennis remains unproven either way at n=113 and
should be judged forward on CLVclosing line valueWhether you got a better price than the market settled at. Widely used as a faster signal of skill than profit, which takes ages to measure., not shut off pre-emptively.
Changes made
Disabled in ev_poll_config (DB-level; INSERT OR IGNORE seeding means the
hardcoded defaults in config.py will not resurrect them):
nba/totalsnhl/totalsncaab/spreadsncaab/totals
Still enabled: mlb/totals, nba/h2h, nba/spreads, ncaab/h2h, nhl/h2h,
nhl/spreads, tennis_atp/h2h, tennis_wta/h2h.
Timing matters: NBA, NCAAB and NHL seasons all resume within ~6–8 weeks, so these markets would otherwise have started firing again.
Open questions
- MLB totals is now the only enabled MLB market, at −8.10% [−17.7, +1.5] over 418 bets. Not proven negative, but concentrating the pipeline into it was an incidental outcome of the August budget rework rather than a decision made on its record. Worth revisiting.
nba/spreadsat −15.51% [−31.2, +0.2] over 150 — CI just barely includes zero, so it was left enabled. Watch it.- ⚠ CORRECTED: the pipeline CAN prove itself on ROI, within about one season. An earlier draft claimed resolving an edge would "take years" — that extrapolated from the seasonal trough. In-season reality: Feb–May ran 11,882 alerts over 4 months (~2,970/month) with only three sports. A Sep–Apr season at that rate is ~23,800 alerts, which clears the 19,232 needed to resolve a 2% edge in under one season; a 3% edge resolves in ~0.4 of one. CLV still moves faster and remains the better in-season signal, but ROI is no longer out of reach on the timescales this project cares about.
- ⚠ The 4th sport is not actually switched on. All three NFL markets
(
nfl/h2h,nfl/spreads,nfl/totals) areenabled=0, and NFL has produced zero alerts in the entire dataset. The Sep–Apr window is currently three sports, not four. Enabling NFL is a deliberate decision that has not been made — worth making consciously before the season rather than by omission. - The market disabling above matters MORE, not less, because of this. NBA totals lost 117.9u across a partial season at 525 settled bets. Those markets were due to resume firing at full volume from October.
Addendum: Pre-Season Hardening Review (2026-08-24)
A second analysis pass substantially corrected the market-family framing above. Recorded here because the correction matters more than the original claim.
⚠ Correction: "totals are broken in every sport" was wrong
The earlier claim that point-based markets are systemically broken conflated alert sources. Splitting totals by source:
| slice | n | ROI | 95% CI | units |
|---|---|---|---|---|
| Model-source totals | 688 | −24.68% | [−31.60, −17.75] | −169.8 |
| Consensus-source totals | 604 | −6.98% | [−14.99, +1.04] | −42.1 |
Every NBA totals alert (525) came from nba_model_v2; every NHL totals alert
(163) from nhl_model_v1. ~69% of the entire −246u came from two model
sources whose totals probabilities were anti-calibrated — stored probability
0.8+ won 2.6% of the time (n=39). Both were root-caused in June 2026 as an
inverted totals mapping, and both are already dead (nba_model_v2 suppressed,
nhl_model_v1 retired).
Consensus-path totals — the path ev_poll_config actually gates — has a CI
that includes zero. It is not established as broken.
Consequence for the disables made above: nba/totals and nhl/totals in
ev_poll_config gate the consensus path, which was roughly neutral, so those
two disables were aimed at the wrong actor and are a mild over-correction. They
are being left off anyway pending the point-match fix below, but on a weaker
justification than originally stated. ncaab/totals and mlb/totals losses
are consensus-path, so those were correctly targeted.
The point-matching hypothesis: partially confirmed
_compare_to_fair (ev_finder.py:914-940) falls back to matching a book's line
against Pinnacle's within |point_diff| ≤ 0.5, and makes no probability
adjustment for the difference. That tolerance only exists since commit
22a41f3 (2026-04-18); before it there was no point check at all, with observed
gaps up to 3.5 points.
Rejoining settled point-market alerts to the at-alert odds snapshots
(unified_odds.db, 8.1M rows) gives the phantom-edge signature directly:
| slice | n | ROI |
|---|---|---|
| MLB totals, exact line match | 19 | +16.0% |
| MLB totals, 0.5-run mismatch | 400 | −6.7% |
| Spreads, book point worse than Pinnacle | 781 | −9.9% |
| Spreads, book point better than Pinnacle | 103 | +31.2% |
The whole live MLB-totals bleed is the mismatch bucket. ±0.5 is not uniformly small: it is ~1pp of win probability in NBA points but 3–8pp in MLB runs / NHL goals, and lands on 3 and 7 in NFL — larger than the 2–3% edges being alerted.
Ruled out as alternative explanations: settlement/grading (re-graded all 1,327
settled totals from final_score, zero discrepancies), devigdevigRemoving the bookmaker's cut from odds to recover the market's actual implied probability. method (methods
agree to <0.3pp near even money), and side inversion on the consensus path.
Ranked hardening actions
- Exact-line matching, or a point-adjusted fair probability. The only verified-live, negative-signed, causally-clean defect. Honest cost: only ~4% of MLB totals alerts had exactly matching lines, so exact-match cuts that market ~95% — correctly, since the volume was phantom. The better version is a per-sport half-point value table that adjusts the fair prob and keeps the alert only if the adjusted edge still clears threshold.
- Credit circuit-breaker. Current config projects 25–35k credits against a 20k cap. August burned ~7.7k (39% of budget) polling NHL/NBA/NFL with zero games scheduled. Also alert on zero billed credits for 24h — that is the signature of the dead-key failure behind the 19-day August outage.
- Blacklist sharp/offshore books as alert TARGETS (betonlineag, lowvig, betus, williamhill_us) while keeping them as consensus inputs: −60.7u across only 324 alerts. When a sharp book diverges from Pinnacle it is leading, not lagging.
- CLV tripwire.
clv>0rows ran +1.15%,clv≤0ran −6.28% (n=2,158 reliable), and the current mean is ≈ −0.15 — the pipeline captures roughly zero edge pre-vig. Weekly per-(sport, market, source) mean CLV; two consecutive weeks below −0.5% auto-disables the market. This detects a broken market in ~10 days instead of six months. - Skip polling sports with no games in the window — frees the budget for the above.
What not to do
- Don't re-enable props for alerts. 8,669 settled at −0.75% while consuming half the credit budget.
- Don't resurrect
nba_model_v2on its +12.9% spreads / +12.2% ML. Those rows are from the fail-open era whenpinnacle_fairwas NULLnull resultA test that found nothing. "Null" is the starting assumption that there is no real effect; a "null result" means the data gave us no reason to abandon that assumption. It does not mean the data was missing or the test failed to run. on nearly all model rows, fabricating a ~13pt edge. Unaudited survivorship of a broken pipeline — the same pattern that produced the "+5.98% methodology v2" backtest that live results refuted. - Don't scale tennis on +14.9% (n=113, CI spans zero) — the pattern that burned 12-1 momentum and NHL v3's +4.32%.
- Don't chase the +2.1% in the 10–20% EV bucket. corr(EV, P&L) = −0.043.
- Don't spend the pre-season on new alpha sources. The 2026-08 cycle already killed Kalshi, arbs, boosts and composites. The measured problem is execution and measurement quality.
Revised bottom line
The six-month record is not evidence that line-shopping against Pinnacle is a −2% strategy. It is evidence that two inverted model-totals paths (now dead), an unadjusted line-matcher, and four sharp books that cannot be beaten were allowed to run for months with no smoke detector. The 246u went unnoticed because evaluation ran on ROI, which needs thousands of bets, instead of CLV, which needs dozens.