New to these reports? Start here
  • Dotted-underlined words have a plain-English definition — hover or tap them. Every term is also on the glossary page.
  • "Null" means we found nothing, not that something broke. Most reports here are negative results, on purpose — knowing an idea doesn't work is the point.
  • Two questions get asked separately. First, is the effect real? Second, is it already priced into the betting odds? An effect can be completely real and still useless to bet on.
  • A "calibration" row is a self-check. It runs the same method on something already known to be true. If that fails, the whole report is unreliable — so it's reported alongside the findings.
  • If a confidence interval includes zero, the real effect might be nothing at all, so no claim gets made.

EV Alert Pipeline — Record Review, August 2026

Review of the live EV alert record and the config changes made as a result. Data: [redacted], 12,821 alerts (Feb–Aug 2026), 12,718 settled win/loss. This is the deduped store (write-time dedup added 2026-07; history collapsed 34.7k → 12.6k), so rows are one-per-pick, not duplicate-weighted.

ROIreturn on investmentProfit as a percentage of the money wagered. +2% means $2 profit per $100 bet. is flat 1u at the actual alert odds, pushes excluded.

Headline

Overall ROI −1.94% [−3.62, −0.26]
Settled alerts 12,718
Net −246.5 units

A small but statistically detectable bleed — the CI excludes zero.

⚠ Recent months are uninformative — but the reason is SEASONALITY, not the budget rework. An earlier draft of this review attributed the volume drop to the August rework. That was wrong. Volume by month and sport:

Month MLB NBA NCAAB NHL Total
2026-02 – 2,042 839 7 2,888
2026-03 – 4,395 559 538 5,492
2026-04 – 594 – 103 697
2026-05 52 2,152 – 601 2,805
2026-06 329 168 – 60 722
2026-07 98 – – – 98
2026-08 107 – – – 119

July–August is MLB-only because MLB is the only sport playing. NBA, NCAAB and NHL all ended. The pipeline is in its seasonal trough, not a degraded state. August's −15% on n=109 (CI spanning 36 points) still means nothing, but "volume collapsed" was the wrong diagnosis.

Where the loss actually is

Four markets accounted for 234 of the 246 units lost, and all four were still enabled at the time of this review:

Market n ROI 95% CI95% confidence intervalThe range the true value is plausibly in. If this range includes zero, we cannot rule out that the real effect is nothing at all. Units
NBA totals 525 −22.46% [−30.5, −14.4] −117.9
NHL totals 163 −31.83% [−45.4, −18.3] −51.9
NCAAB spreads 592 −8.67% [−16.7, −0.7] −51.3
NCAAB totals 48 −26.80% [−53.5, −0.1] −12.9

Counterfactual with those four removed: −1.94% → −0.11%, net −246.5u → −12.5u. Everything else is roughly break-even.

Why this is not just post-hoc subgroup mining

Selecting losers by inspecting their record is exactly the grid-argmax trap this project has documented as skill-less, so it deserves scepticism. Two things override it here:

  1. The June 2026 ROI audit already flagged "NBA totals as the big leaks." This is a replication of a previously-identified finding that was never actioned — not a fresh discovery.
  2. NBA totals at −22.5% with an upper CI bound of −14.4% over 525 bets is far outside what chance produces across ~11 markets tested.

Confidence is not uniform: NBA/NHL totals are strong; NCAAB spreads is moderate; NCAAB totals (n=48, CI upper bound −0.07%) is marginal and was disabled mainly because its volume is tiny and the cost of being wrong is near zero.

The EV number itself is not working

EV bucket n ROI 95% CI
0–2% 5,559 −1.56% [−3.99, +0.87]
2–4% 4,178 −1.95% [−4.91, +1.02]
4–6% 1,305 +1.24% [−4.18, +6.66]
6–10% 564 +0.14% [−8.29, +8.57]
10–20% 371 +2.10% [−8.26, +12.46]
20%+ 696 −15.67% [−23.57, −7.76]

corr(predicted EV, realised P&L) = −0.043 across 12,718 alerts. Predicted EV carries essentially no information about realised profit.

⚠ Correction: the EV cap already exists and already works

The initial recommendation from this review was "deploy an EV cap." That was wrong — EVFinder.MAX_EV_PCT = 8.0 is in the code and is applied. Checking the data confirms it works: the last market_consensus alert above 8% EV was in February 2026.

Decomposing the 20%+ bucket by source shows the penalty is not a general law but two specific model sources:

Source at EV ≥ 20% n ROI 95% CI
nba_model_v2 530 −17.96% [−26.41, −9.50]
nhl_model_v1 71 −28.17% [−47.74, −8.60]
tennis_elo_v1 94 +4.24% [−25.56, +34.05]

Both offenders are already handled: nba_model_v2 is capped at 4.0 EV (alert_tracker.py), and nhl_model_v1 is retired. No further EV cap was added.

Specifically, tennis_elo_v1 was left uncapped despite emitting 28–35% EV alerts, because its high-EV subset is the one that is not losing (+4.24%) and tennis overall is the only positive segment (+14.87%, n=113). Capping it would have been acting against the evidence in order to satisfy a rule derived from a different, already-fixed source. Tennis remains unproven either way at n=113 and should be judged forward on CLVclosing line valueWhether you got a better price than the market settled at. Widely used as a faster signal of skill than profit, which takes ages to measure., not shut off pre-emptively.

Changes made

Disabled in ev_poll_config (DB-level; INSERT OR IGNORE seeding means the hardcoded defaults in config.py will not resurrect them):

  • nba/totals
  • nhl/totals
  • ncaab/spreads
  • ncaab/totals

Still enabled: mlb/totals, nba/h2h, nba/spreads, ncaab/h2h, nhl/h2h, nhl/spreads, tennis_atp/h2h, tennis_wta/h2h.

Timing matters: NBA, NCAAB and NHL seasons all resume within ~6–8 weeks, so these markets would otherwise have started firing again.

Open questions

  • MLB totals is now the only enabled MLB market, at −8.10% [−17.7, +1.5] over 418 bets. Not proven negative, but concentrating the pipeline into it was an incidental outcome of the August budget rework rather than a decision made on its record. Worth revisiting.
  • nba/spreads at −15.51% [−31.2, +0.2] over 150 — CI just barely includes zero, so it was left enabled. Watch it.
  • ⚠ CORRECTED: the pipeline CAN prove itself on ROI, within about one season. An earlier draft claimed resolving an edge would "take years" — that extrapolated from the seasonal trough. In-season reality: Feb–May ran 11,882 alerts over 4 months (~2,970/month) with only three sports. A Sep–Apr season at that rate is ~23,800 alerts, which clears the 19,232 needed to resolve a 2% edge in under one season; a 3% edge resolves in ~0.4 of one. CLV still moves faster and remains the better in-season signal, but ROI is no longer out of reach on the timescales this project cares about.
  • ⚠ The 4th sport is not actually switched on. All three NFL markets (nfl/h2h, nfl/spreads, nfl/totals) are enabled=0, and NFL has produced zero alerts in the entire dataset. The Sep–Apr window is currently three sports, not four. Enabling NFL is a deliberate decision that has not been made — worth making consciously before the season rather than by omission.
  • The market disabling above matters MORE, not less, because of this. NBA totals lost 117.9u across a partial season at 525 settled bets. Those markets were due to resume firing at full volume from October.

Addendum: Pre-Season Hardening Review (2026-08-24)

A second analysis pass substantially corrected the market-family framing above. Recorded here because the correction matters more than the original claim.

⚠ Correction: "totals are broken in every sport" was wrong

The earlier claim that point-based markets are systemically broken conflated alert sources. Splitting totals by source:

slice n ROI 95% CI units
Model-source totals 688 −24.68% [−31.60, −17.75] −169.8
Consensus-source totals 604 −6.98% [−14.99, +1.04] −42.1

Every NBA totals alert (525) came from nba_model_v2; every NHL totals alert (163) from nhl_model_v1. ~69% of the entire −246u came from two model sources whose totals probabilities were anti-calibrated — stored probability 0.8+ won 2.6% of the time (n=39). Both were root-caused in June 2026 as an inverted totals mapping, and both are already dead (nba_model_v2 suppressed, nhl_model_v1 retired).

Consensus-path totals — the path ev_poll_config actually gates — has a CI that includes zero. It is not established as broken.

Consequence for the disables made above: nba/totals and nhl/totals in ev_poll_config gate the consensus path, which was roughly neutral, so those two disables were aimed at the wrong actor and are a mild over-correction. They are being left off anyway pending the point-match fix below, but on a weaker justification than originally stated. ncaab/totals and mlb/totals losses are consensus-path, so those were correctly targeted.

The point-matching hypothesis: partially confirmed

_compare_to_fair (ev_finder.py:914-940) falls back to matching a book's line against Pinnacle's within |point_diff| ≤ 0.5, and makes no probability adjustment for the difference. That tolerance only exists since commit 22a41f3 (2026-04-18); before it there was no point check at all, with observed gaps up to 3.5 points.

Rejoining settled point-market alerts to the at-alert odds snapshots (unified_odds.db, 8.1M rows) gives the phantom-edge signature directly:

slice n ROI
MLB totals, exact line match 19 +16.0%
MLB totals, 0.5-run mismatch 400 −6.7%
Spreads, book point worse than Pinnacle 781 −9.9%
Spreads, book point better than Pinnacle 103 +31.2%

The whole live MLB-totals bleed is the mismatch bucket. ±0.5 is not uniformly small: it is ~1pp of win probability in NBA points but 3–8pp in MLB runs / NHL goals, and lands on 3 and 7 in NFL — larger than the 2–3% edges being alerted.

Ruled out as alternative explanations: settlement/grading (re-graded all 1,327 settled totals from final_score, zero discrepancies), devigdevigRemoving the bookmaker's cut from odds to recover the market's actual implied probability. method (methods agree to <0.3pp near even money), and side inversion on the consensus path.

Ranked hardening actions

  1. Exact-line matching, or a point-adjusted fair probability. The only verified-live, negative-signed, causally-clean defect. Honest cost: only ~4% of MLB totals alerts had exactly matching lines, so exact-match cuts that market ~95% — correctly, since the volume was phantom. The better version is a per-sport half-point value table that adjusts the fair prob and keeps the alert only if the adjusted edge still clears threshold.
  2. Credit circuit-breaker. Current config projects 25–35k credits against a 20k cap. August burned ~7.7k (39% of budget) polling NHL/NBA/NFL with zero games scheduled. Also alert on zero billed credits for 24h — that is the signature of the dead-key failure behind the 19-day August outage.
  3. Blacklist sharp/offshore books as alert TARGETS (betonlineag, lowvig, betus, williamhill_us) while keeping them as consensus inputs: −60.7u across only 324 alerts. When a sharp book diverges from Pinnacle it is leading, not lagging.
  4. CLV tripwire. clv>0 rows ran +1.15%, clv≤0 ran −6.28% (n=2,158 reliable), and the current mean is ≈ −0.15 — the pipeline captures roughly zero edge pre-vig. Weekly per-(sport, market, source) mean CLV; two consecutive weeks below −0.5% auto-disables the market. This detects a broken market in ~10 days instead of six months.
  5. Skip polling sports with no games in the window — frees the budget for the above.

What not to do

  • Don't re-enable props for alerts. 8,669 settled at −0.75% while consuming half the credit budget.
  • Don't resurrect nba_model_v2 on its +12.9% spreads / +12.2% ML. Those rows are from the fail-open era when pinnacle_fair was NULLnull resultA test that found nothing. "Null" is the starting assumption that there is no real effect; a "null result" means the data gave us no reason to abandon that assumption. It does not mean the data was missing or the test failed to run. on nearly all model rows, fabricating a ~13pt edge. Unaudited survivorship of a broken pipeline — the same pattern that produced the "+5.98% methodology v2" backtest that live results refuted.
  • Don't scale tennis on +14.9% (n=113, CI spans zero) — the pattern that burned 12-1 momentum and NHL v3's +4.32%.
  • Don't chase the +2.1% in the 10–20% EV bucket. corr(EV, P&L) = −0.043.
  • Don't spend the pre-season on new alpha sources. The 2026-08 cycle already killed Kalshi, arbs, boosts and composites. The measured problem is execution and measurement quality.

Revised bottom line

The six-month record is not evidence that line-shopping against Pinnacle is a −2% strategy. It is evidence that two inverted model-totals paths (now dead), an unadjusted line-matcher, and four sharp books that cannot be beaten were allowed to run for months with no smoke detector. The 246u went unnoticed because evaluation ran on ROI, which needs thousands of bets, instead of CLV, which needs dozens.

On this page

Terms in this report

Source

backtests/ev_record_2026_08/FINDINGS.md
updated 2026-08-24 00:56