New to these reports? Start here
  • Dotted-underlined words have a plain-English definition — hover or tap them. Every term is also on the glossary page.
  • "Null" means we found nothing, not that something broke. Most reports here are negative results, on purpose — knowing an idea doesn't work is the point.
  • Two questions get asked separately. First, is the effect real? Second, is it already priced into the betting odds? An effect can be completely real and still useless to bet on.
  • A "calibration" row is a self-check. It runs the same method on something already known to be true. If that fails, the whole report is unreliable — so it's reported alongside the findings.
  • If a confidence interval includes zero, the real effect might be nothing at all, so no claim gets made.

Pre-registration: per-market CLV judging, 2026-27 season

Written 2026-08-24. Locked before the season starts. Status: ACTIVE — no data has been observed under this rule.

Why this document exists

This project keeps finding edges that evaporate. The Retrosheet split miner found 96.6% of naive p<0.05 "trends" died out-of-sampleout-of-sampleTested on data that was not used to build or tune the idea. This is the honest test; results on the data you built with are almost always flattering.. The NHL v3 thresholds turned out to be grid-argmax with no skill. The OSINT sweep combined 50 zero signals into an in-samplein-sampleMeasured on the same data used to build or tune the idea. Nearly always looks better than reality. SharpeSharpe ratioReturn relative to how much it bounced around. Higher means smoother returns for the same profit. of 1.80 while out-of-sample stayed at 0.00.

And most recently: the four markets disabled in the August 2026 record review were discovered and scored in the same analysis. So was the sharp-book blacklist. Those may well be real — nba/totals had been independently flagged in June — but their measured improvement is not trustworthy, because nothing separated the finding from the fitting.

The purpose of this document is to make the 2026-27 season a genuine test. The segment list, the metric, the sample milestones, the error control, and the actions are all fixed now, while there is no data to fit them to. If the rule is embarrassing in February, that is the rule working.

Primary metric

Mean clv_fair_pct per segment, over eligible observations.

clv_fair_pct is closing-line value measured against Pinnacle's two-sided devigged fair probability. It is the right metric and ROIreturn on investmentProfit as a percentage of the money wagered. +2% means $2 profit per $100 bet. is not, for a reason that is now quantified: at observed dispersion a 1-point CLVclosing line valueWhether you got a better price than the market settled at. Widely used as a faster signal of skill than profit, which takes ages to measure. effect resolves in ~30-40 observations, while a 2% ROI edge needs ~17,500 bets. At the current config's in-season rate (~455 alerts/month, since props are disabled and were 82% of historical volume) ROI would take 4-5 seasons. CLV takes weeks.

ROI is recorded but is NOT a decision input this season. It cannot resolve in the time available, and having two metrics with different speeds invites picking whichever agrees.

Eligibility

An observation counts if and only if: - clv_fair_pct IS NOT NULL and clv_unreliable = 0 - the book is not in SOFT_BOOK_BLOCKLIST - the source is live (not a retired model) - the market is enabled in ev_poll_config

Segments and baseline parameters

Measured on Feb-Aug 2026 under the current config. These are the priors, not the test — the test uses data generated after 2026-09-01.

segment alerts CLV n coverage mean CLV sd judgeable
ncaab/h2h 565 541 96% −0.30% 2.88 YES
nba/h2h 192 186 97% −0.22% 3.31 YES
tennis_wta/h2h 110 98 89% +0.95% 3.32 YES
mlb/h2h 38 36 95% −0.16% 2.79 YES
nhl/spreads 25 22 88% −0.22% 1.66 YES
mlb/spreads 21 19 90% −0.75% 2.06 YES
nhl/h2h 10 9 90% +1.37% 1.69 YES
tennis_atp/h2h 5 5 100% +1.84% 1.52 YES
nfl/h2h 0 0 — — — YES (no prior)
mlb/totals 216 84 39% −0.91% 1.41 BLOCKED
nba/spreads 150 27 18% −1.26% 0.93 BLOCKED

NFL has never fired an alert in this system's history (enabled=0 until 2026-08). It has no prior at all. September is its first-ever data and the first test of whether NFL CLV capture even works.

Blocked segments — must be unblocked before they may be judged

mlb/totals and nba/spreads are excluded from decisions until two things are fixed and re-measured:

  1. The CLV line-match bug, which contaminates all spread/total CLV.
  2. Capture coverage. These are the only segments below 88% coverage (39% and 18%). They also show anomalously compressed dispersion — nba/spreads sd 0.93 against nba/h2h sd 3.31 for the same sport. Low coverage plus low variance is the signature of a selection effect in the capture itself: we are likely only capturing closes on the non-random subset where line matching succeeded. A mean computed on that subset is biased and no sample size fixes it.

Unblock criterion: coverage ≥85% AND dispersion within a factor of 2 of the same sport's h2h segment. Until then these segments keep alerting and keep recording, but generate no disable/keep decision.

Decision procedure

Milestones. Each segment is judged when its eligible-observation count first crosses n = 150, then n = 300, then n = 600. Not weekly, not continuously — at these three points only. This is what controls the number of looks.

Why not weekly: ~8 live segments checked weekly from November to February is ~136 tests. At a naive α=0.05 that produces ~7 false disables by chance alone, which would silently dismantle the pipeline. Three milestones across ~8 segments is ~24 tests.

Test. At each milestone, for every segment that has reached it: - H0: mean CLV = 0, two-sided, one-sample t-test - Collect p-values across all segments judged at that milestone - Apply Benjamini-Hochberg at FDRfalse discovery rateA method for handling many simultaneous tests, controlling what share of your "discoveries" are expected to be flukes. = 0.10 across that set

Actions:

BH-significant? sign action
yes negative DISABLE the market in ev_poll_config
yes positive KEEP; eligible for size increase
no either CONTINUE — no action, collect to next milestone

A segment that reaches n=600 without significance is declared NULLnull resultA test that found nothing. "Null" is the starting assumption that there is no real effect; a "null result" means the data gave us no reason to abandon that assumption. It does not mean the data was missing or the test failed to run.: it is kept enabled at baseline size but is no longer a candidate for scaling, and it is not judged again this season.

Re-test policy. A disabled segment is re-enabled and re-tested at the start of the following season. Disables are not permanent. Without this the rule ratchets monotonically toward zero markets, since noise can only ever remove.

Calendar

NCAAB is the volume engine (~280 alerts/month, ~60% of in-season volume) and does not start until November. Therefore:

  • September — shakedown, not judgement. Confirm NFL fires at all and its CLV captures. Collect tennis US Open data on the one positive segment. Fix the line-match bug and audit coverage. No decisions.
  • October — NFL + MLB playoffs; NBA/NHL start late. First milestones possible for tennis and NFL.
  • November onward — all four sports. ncaab/h2h reaches n=150 in roughly three weeks; nba/h2h in about two months.

Predictions (recorded so they can be wrong)

  1. Overall mainline CLV finishes negative but small, in [−0.5%, 0%].
  2. tennis is the segment most likely to be significantly positive. If any segment survives at FDR 0.10 with a positive sign, tennis is the prior favourite.
  3. NFL will produce far fewer alerts than expected — it is a low-count, heavily-priced market and this system has never tested it.
  4. At least one segment currently negative will regress toward zero, because part of the August improvement was post-hoc.

Explicitly forbidden this season

  • Tuning the threshold, the FDR level, or the milestones after seeing data.
  • Judging any segment not in the table above, or defining new segments by splitting on book, time-of-day, EV bucket, or line movement. Post-hoc subgroups are how the four disabled markets were found, and their measured benefit is not trustworthy for exactly that reason.
  • Building a composite score across weak signals. Prior result: combining 50 zero-signals lifted in-sample Sharpe 0.34 to 1.80 while OOS stayed at 0.00.
  • Using ROI to overturn a CLV decision.

Amendments

Any change to this document must be appended below with a date and a reason, and must state whether data had been observed at the time. A change made after seeing data invalidates the pre-registration for the affected segments — say so plainly rather than editing above.

Amendment 1 — 2026-08-24 — eligibility rule for point markets

Data observed at time of amendment: NONE. The season has not started; no observation exists under this rule. This amendment is therefore made blind and does not compromise the pre-registration.

What was wrong. The eligibility criterion above requires clv_unreliable = 0. Investigation the same day (backtests/clv_line_adjustment/FINDINGS.md) established that this flag is exactly a line-movement flag: rows with clv_unreliable = 0 matched the closing lineclosing lineThe final odds right before a game starts. It reflects everything the betting market knows, which makes it the hardest benchmark to beat. exactly 100.0% of the time, rows with clv_unreliable = 1 matched 0.0% of the time. The criterion therefore does not exclude bad measurements — it excludes every game whose line moved, which is 70% of spread/total alerts. That is missing-not-at-random: it halves the variance (sd 3.99 → 1.84) and pulls the mean toward zero (−0.35% → −0.20%).

Left unamended, this rule would have judged every point market on the subset of games where nothing happened, and would have done so with an artificially narrow spread — inflating apparent significance.

Amended eligibility, point markets (spreads, totals) only:

An observation is eligible if the closing price is taken at the alert's exact line at the closing snapshot, from any book — not Pinnacle alone. Measured recovery: 44.8% of currently-flagged alerts have some book still quoting the exact line at close. The book quality tier must be recorded (Pinnacle vs soft), because a soft book's close is a worse probability estimate, and every milestone report must state the result both pooled and restricted to Pinnacle.

Alerts whose line is gone at every book by close (55.2%) remain excluded and counted, pending the point-pricing work. The excluded count must be reported at each milestone so the size of the remaining hole stays visible.

Moneyline (h2h) segments are unaffected — there is no line to move, which is why their capture coverage is already 88–100%. The headline CLV numbers in this document's segment table are h2h-dominated and stand.

Consequence for the blocked segments. mlb/totals and nba/spreads remain BLOCKED. Their unblock criterion is unchanged (coverage ≥85%, dispersion within 2x the same sport's h2h) but is now expected to be met by fixing capture rather than by waiting for volume.

Not amended: milestones, FDR level, actions, re-test policy, forbidden analyses. Only the eligibility definition changed, and only for point markets.

Amendment 2 — 2026-08-24 — capture tiers, and the blocked segments unblock

Data observed at time of amendment: NONE. Season not started. Made blind.

Amendment 1 widened eligibility to any book quoting the exact line at close, recovering 44.8% of the discarded rows. The remaining rows have now been made measurable by pricing the point difference against the empirical distribution of the result versus the closing line (backtests/clv_line_adjustment/). Validated on 338 alerts whose true exact-line close was independently recoverable: MAEmean absolute errorAverage size of the miss, ignoring direction. If a projection is off by 3 one week and -5 the next, the MAE is 4. Lower is better. 4.47pp → 1.22pp, biasbiasWhether the misses lean consistently one way. A projection can have a good average error size but still be biased if it is almost always too high. Bias is often the more fixable problem. ~0.

Amended eligibility, point markets. An observation is eligible in any of three tiers, and every milestone report must break results out by tier:

tier definition
T1 exact our line was still the closing line
T2 recovered another book quoted our exact line at the close
T3 adjusted no book held it; the difference was priced (~1.2pp added noise)

Rows that remain unpriceable stay excluded and counted. NCAAB point markets have no residual table and are not pinned-line, so they can only ever reach T1/T2 — both are disabled markets, so this blocks nothing.

mlb/totals and nba/spreads are UNBLOCKED. Their criterion — coverage ≥85% and dispersion within 2× the same sport's h2h — was fixed in advance and is a pure measurement-quality test, not an outcome test, so applying it now does not peek. Both reach 100% coverage, at dispersion ratios 1.14 and 1.58.

⚠ Expect point markets to look worse than the segment table above. That table was built on T1 only, which is the subset where nothing happened. Historically the tiers run −0.20% / −1.04% / −2.43%, pooling to −1.08% against the −0.20% previously reported. The priors in this document's segment table are therefore optimistic for every spread/total row and should not be read as predictions. Prediction 1 (overall mainline CLV in [−0.5%, 0%]) applies to the h2h-dominated aggregate and is left standing as written.

Not amended: milestones, FDR level, actions, re-test policy, forbidden analyses, predictions.

On this page

Terms in this report

Source

backtests/ev_clv_prereg_2026_09/PREREGISTRATION.md
updated 2026-08-24 23:38