New to these reports? Start here
  • Dotted-underlined words have a plain-English definition — hover or tap them. Every term is also on the glossary page.
  • "Null" means we found nothing, not that something broke. Most reports here are negative results, on purpose — knowing an idea doesn't work is the point.
  • Two questions get asked separately. First, is the effect real? Second, is it already priced into the betting odds? An effect can be completely real and still useless to bet on.
  • A "calibration" row is a self-check. It runs the same method on something already known to be true. If that fails, the whole report is unreliable — so it's reported alongside the findings.
  • If a confidence interval includes zero, the real effect might be nothing at all, so no claim gets made.

The stable movement filter is unvalidated — and unfalsifiable as deployed

2026-08-28. ev_finder.py silently discards every opportunity classified movement_class == 'stable' before it can alert. It has done so since ~13 Feb 2026, on the strength of one number: −17.76% ROIreturn on investmentProfit as a percentage of the money wagered. +2% means $2 profit per $100 bet., N=411, p=0.0005.

It drops 53% of enriched opportunities (491 of 923 in the only window where the data exists). That is a large, permanent, unreviewed reduction in what the pipeline is allowed to bet.

The filter deletes its own evidence

stable rows exist in ev_opportunities only between 2026-02-01 and 2026-02-13. Every sibling class spans February to August. The reason is structural: from the day the filter shipped, suppressed opportunities were never saved, so no stable row has been recorded since.

That makes the rule unfalsifiable in production. Any re-test is run on the same sample that produced it — there is no out-of-sampleout-of-sampleTested on data that was not used to build or tune the idea. This is the honest test; results on the data you built with are almost always flattering. data and, under the current design, there never can be.

Within its own window the effect looks real

Restricting every class to 2026-02-01..02-13, so the comparison is like-for-like:

class n ROI vs rest t
stable 454 −20.23% −1.20% −2.73
expanding 208 +10.03% −18.11% +3.38
confirming 115 −17.08% −10.38% −0.67
fading 80 −7.60% −11.66% +0.35

Taken alone, stable at t = −2.73 looks like a finding.

But the window does not predict the future

The classes that can be checked out of sample both reversed:

class Feb 1–13 after Feb 13
expanding +10.03% −3.31%
confirming −17.08% +4.52%
fading −7.60% +1.07%

expanding had a STRONGER in-window t-statistic than stable (+3.38 vs −2.73) and its effect vanished — and reversed — afterwards. confirming scored −17.08% in that window, close to stable's −20.23%; a filter built on the same reasoning would have killed it, and it has run +4.52% since.

A 13-day window plainly does not identify a persistent effect in this data. The stable result is not distinguishable from the same artefact, and it is the one class we cannot check because the filter stopped recording it.

Verdict

Not proven, not disproven — and that is the problem. The evidence for discarding half the pipeline's enriched opportunities is one 13-day slice whose siblings all reversed out of sample.

Deliberately not removed. Removing it would be acting on the same weak evidence in the opposite direction, and if stable really is −20% that costs money immediately. Instead the filter now runs in shadow mode: still suppressed from alerting, so live behaviour and the track record are unchanged, but recorded, so it can finally be judged on data it did not generate.

Verdict conditions, set now

Judge on stable rows found after 2026-08-28, against the contemporaneous non-stable rows:

  • Remove the filter if stable's ROI is within ~5pp of the others, or its CI spans theirs, at n ≥ 300.
  • Keep it if stable is worse by >10pp with a CI excluding zero at n ≥ 300.
  • Otherwise keep collecting. Do not judge before n = 300; that is roughly the sample that produced the original claim, and the whole lesson here is that it was not enough.

Supporting change

Opportunity settlement previously piggybacked on ALERT settlement, so a suppressed opportunity could never receive an outcome — the reason this went unchecked for six months. scripts/odds/settle_shadow_opportunities.py now grades opportunities directly from unified_odds.db::game_results.

⚠ Props cannot be graded this way (they need player stats, not scores). All prop markets are disabled, so this does not affect the experiment.

On this page

Terms in this report

Source

backtests/stable_filter_audit/FINDINGS.md
updated 2026-08-29 02:22