New to these reports? Start here
  • Dotted-underlined words have a plain-English definition — hover or tap them. Every term is also on the glossary page.
  • "Null" means we found nothing, not that something broke. Most reports here are negative results, on purpose — knowing an idea doesn't work is the point.
  • Two questions get asked separately. First, is the effect real? Second, is it already priced into the betting odds? An effect can be completely real and still useless to bet on.
  • A "calibration" row is a self-check. It runs the same method on something already known to be true. If that fails, the whole report is unreliable — so it's reported alongside the findings.
  • If a confidence interval includes zero, the real effect might be nothing at all, so no claim gets made.

Teaser/parlay legs are independent — the correlation worry was unfounded

2026-08-28. The teaser EV maths multiplies per-leg probabilities, which assumes the legs are independent. That assumption was flagged as unmeasured and it drives a live decision (2-team vs 3-team teaser). Measured it on 11 seasons of NFL schedules (2015–2025, 3,028 games).

Result: no detectable correlation

Wong dog legs (underdog +1.5 to +2.5, teased 6 points), grouped by season-week, all combinations within each week, bootstrapped clustered by week:

legs observed P(all win) if independent 95% CI95% confidence intervalThe range the true value is plausibly in. If this range includes zero, we cannot rule out that the real effect is nothing at all. verdict
2 0.6115 0.5972 [0.5427, 0.6813] consistent with independence
3 0.4545 0.4616 [0.3551, 0.5772] consistent with independence
4 0.2978 0.3567 [0.1818, 0.4755] consistent with independence

The better-powered general test — every game's dog-ATS outcome, 20,623 within-week pairs — is tighter and also nullnull resultA test that found nothing. "Null" is the starting assumption that there is no real effect; a "null result" means the data gave us no reason to abandon that assumption. It does not mean the data was missing or the test failed to run.: observed 0.2504 against 0.2533 independent, a difference of −0.0029.

So there is no week-level common factor moving NFL sides together at any magnitude that matters. Multiplying leg probabilities is fine. The teaser EV figures stand unchanged, and the earlier conclusion holds: 3-team is the sweet spot, and 4-team collapses on variance and leg supply, not correlationcorrelationHow closely two things move together, from -1 (opposite) through 0 (unrelated) to +1 (in lockstep). It does not by itself mean one causes the other..

The trap this nearly walked into

Before clustering the bootstrapbootstrapRe-running a calculation on thousands of resampled versions of the data to see how much the answer wobbles. The spread of those answers becomes the confidence interval., the k=4 point estimate looked alarming: observed 0.2978 against 0.3567 independent, which would have cut 4-team teaser EV from +24.13% to +3.62%, a 20.5pp swing, and killed the strategy.

It is noise. Only 39 weeks ever contain four qualifying legs, and the 225 "combinations" drawn from them overlap heavily — they are nowhere near 225 independent observations. Treating them as such produces a confident wrong answer. Clustering by week widens the interval to [0.1818, 0.4755], which comfortably contains independence.

Same failure mode as the stable movement filter audited the same day: a small-sample slice that looks decisive until the dependence structure is respected.

What actually drives teaser EV

Not correlation — the leg probability itself:

p_leg P(all 3) EV at 3-team +175
0.720 0.3732 +2.6%
0.740 0.4052 +11.4%
0.756 0.4320 +18.8%
0.776 0.4673 +28.5%
0.800 0.5120 +40.8%

A 0.08 swing in p_leg moves EV ~38pp; correlation moves it ~0. The in-samplein-sampleMeasured on the same data used to build or tune the idea. Nearly always looks better than reality. estimate is 0.7728 and the out-of-sampleout-of-sampleTested on data that was not used to build or tune the idea. This is the honest test; results on the data you built with are almost always flattering. one 0.756, and that gap alone is worth more than every correlation effect measured here.

Conclusion: stop worrying about leg correlation, and spend the effort tightening p_leg instead.

Scope

NFL only. Other sports were not tested — NFL is the only one with the schedule history and the only sport where a teaser edge exists. The null is about same-week correlation across different games; it says nothing about same-game parlays, whose legs are correlated by construction.

On this page

Terms in this report

Source

backtests/leg_correlation/FINDINGS.md
updated 2026-08-29 02:29