New to these reports? Start here
- Dotted-underlined words have a plain-English definition — hover or tap them. Every term is also on the glossary page.
- "Null" means we found nothing, not that something broke. Most reports here are negative results, on purpose — knowing an idea doesn't work is the point.
- Two questions get asked separately. First, is the effect real? Second, is it already priced into the betting odds? An effect can be completely real and still useless to bet on.
- A "calibration" row is a self-check. It runs the same method on something already known to be true. If that fails, the whole report is unreliable — so it's reported alongside the findings.
- If a confidence interval includes zero, the real effect might be nothing at all, so no claim gets made.
The stable movement filter is unvalidated — and unfalsifiable as deployed
2026-08-28. ev_finder.py silently discards every opportunity classified
movement_class == 'stable' before it can alert. It has done so since ~13 Feb
2026, on the strength of one number: −17.76% ROIreturn on investmentProfit as a percentage of the money wagered. +2% means $2 profit per $100 bet., N=411, p=0.0005.
It drops 53% of enriched opportunities (491 of 923 in the only window where the data exists). That is a large, permanent, unreviewed reduction in what the pipeline is allowed to bet.
The filter deletes its own evidence
stable rows exist in ev_opportunities only between 2026-02-01 and
2026-02-13. Every sibling class spans February to August. The reason is
structural: from the day the filter shipped, suppressed opportunities were never
saved, so no stable row has been recorded since.
That makes the rule unfalsifiable in production. Any re-test is run on the same sample that produced it — there is no out-of-sampleout-of-sampleTested on data that was not used to build or tune the idea. This is the honest test; results on the data you built with are almost always flattering. data and, under the current design, there never can be.
Within its own window the effect looks real
Restricting every class to 2026-02-01..02-13, so the comparison is like-for-like:
| class | n | ROI | vs rest | t |
|---|---|---|---|---|
| stable | 454 | −20.23% | −1.20% | −2.73 |
| expanding | 208 | +10.03% | −18.11% | +3.38 |
| confirming | 115 | −17.08% | −10.38% | −0.67 |
| fading | 80 | −7.60% | −11.66% | +0.35 |
Taken alone, stable at t = −2.73 looks like a finding.
But the window does not predict the future
The classes that can be checked out of sample both reversed:
| class | Feb 1–13 | after Feb 13 |
|---|---|---|
| expanding | +10.03% | −3.31% |
| confirming | −17.08% | +4.52% |
| fading | −7.60% | +1.07% |
expanding had a STRONGER in-window t-statistic than stable (+3.38 vs
−2.73) and its effect vanished — and reversed — afterwards. confirming
scored −17.08% in that window, close to stable's −20.23%; a filter built on
the same reasoning would have killed it, and it has run +4.52% since.
A 13-day window plainly does not identify a persistent effect in this data. The
stable result is not distinguishable from the same artefact, and it is the
one class we cannot check because the filter stopped recording it.
Verdict
Not proven, not disproven — and that is the problem. The evidence for discarding half the pipeline's enriched opportunities is one 13-day slice whose siblings all reversed out of sample.
Deliberately not removed. Removing it would be acting on the same weak
evidence in the opposite direction, and if stable really is −20% that costs
money immediately. Instead the filter now runs in shadow mode: still
suppressed from alerting, so live behaviour and the track record are unchanged,
but recorded, so it can finally be judged on data it did not generate.
Verdict conditions, set now
Judge on stable rows found after 2026-08-28, against the contemporaneous
non-stable rows:
- Remove the filter if stable's ROI is within ~5pp of the others, or its CI spans theirs, at n ≥ 300.
- Keep it if stable is worse by >10pp with a CI excluding zero at n ≥ 300.
- Otherwise keep collecting. Do not judge before n = 300; that is roughly the sample that produced the original claim, and the whole lesson here is that it was not enough.
Supporting change
Opportunity settlement previously piggybacked on ALERT settlement, so a
suppressed opportunity could never receive an outcome — the reason this went
unchecked for six months. scripts/odds/settle_shadow_opportunities.py now
grades opportunities directly from unified_odds.db::game_results.
⚠ Props cannot be graded this way (they need player stats, not scores). All prop markets are disabled, so this does not affect the experiment.