New to these reports? Start here
  • Dotted-underlined words have a plain-English definition — hover or tap them. Every term is also on the glossary page.
  • "Null" means we found nothing, not that something broke. Most reports here are negative results, on purpose — knowing an idea doesn't work is the point.
  • Two questions get asked separately. First, is the effect real? Second, is it already priced into the betting odds? An effect can be completely real and still useless to bet on.
  • A "calibration" row is a self-check. It runs the same method on something already known to be true. If that fails, the whole report is unreliable — so it's reported alongside the findings.
  • If a confidence interval includes zero, the real effect might be nothing at all, so no claim gets made.

TD models: not unpredictable, just pointed at the wrong output

Date: 2026-08-30 Harness: backtests/fantasy_eval/td_probability_experiment.py (re-runnable) Verdict: ✅ PARTIAL — ships for rushing and receiving TDs, NOT for passing TDs

The question

The August audit (fantasy model audit and the memory record) found the TD point-projections are worse than predicting zero, and correctly so: week-to-week TD rate barely persists (r = 0.05–0.19). That was read as "TDs are unpredictable."

That conclusion conflated two different things. A point estimate is the wrong output for a rare count — MAEmean absolute errorAverage size of the miss, ignoring direction. If a projection is off by 3 one week and -5 the next, the MAE is 4. Lower is better. against a mostly-zero target is won by a model that always says 0.0. The question a TD model should answer is does this player score? So: does the same projection carry information about that?

It does. The projections rank well and are calibrated badly

Poisson P(≥1) = 1 − exp(−λ) from the existing point projection, scored against actual ≥ 1 on the full graded history (position-filtered):

stat AUC log losslog lossA score for probability forecasts that punishes confident wrong answers harshly. Lower is better. (raw) log loss (base rate)
passing_tds 0.663 0.6847 0.5459
rushing_tds 0.687 0.5927 0.5248
receiving_tds 0.674 0.5684 0.4496

AUC ≈ 0.67 on all three — the projections genuinely rank who scores. But raw log loss is worse than simply quoting the base rate on all three. Good discrimination with bad calibrationcalibrationA deliberate sanity check on the method itself: run it on something already known to be true. If it fails to detect the known thing, the method is broken and its other results mean nothing. is precisely what a calibration map fixes, and the point estimate was discarding the usable half.

Calibration flips them

Walk-forwardwalk-forwardEvaluating week by week using only what was knowable before each week, mimicking how the model would actually have been used at the time. by season (train on all prior seasons, isotonic on the Poisson probability). Gain = per-row log-loss improvement over the base rate, bootstrapped clustered by player — the same player recurs every week, so iid rows would understate the interval.

stat N LL base LL raw LL isotonic LL Platt gain 95% CI95% confidence intervalThe range the true value is plausibly in. If this range includes zero, we cannot rule out that the real effect is nothing at all. seasons better
passing_tds 1,498 0.5557 0.7411 0.5421 0.5492 +0.0136 [−0.0108, +0.0409] 3/4
rushing_tds 4,738 0.5232 0.5979 0.4854 0.4840 +0.0378 [+0.0264, +0.0501] 4/4
receiving_tds 11,097 0.4470 0.5696 0.4204 0.4214 +0.0266 [+0.0207, +0.0332] 4/4

rushing and receiving are established — CI excludes zero, every test season improves. passing_tds is not — its CI spans zero and one season is worse. That is unsurprising: QBs have a 76% base rate (little room) and only ~1.5k graded rows (little power). It is reported as unproven rather than averaged in.

Isotonic and Platt are within noise of each other. Isotonic ships, as it assumes no shape for the Poisson-link mismatch.

⚠️ Never expose the raw Poisson probability

It is worse than climatology on all three stats — passing_tds blows up to 1.3375 in 2025 alone. td_probability.p_at_least_one() returns None when it cannot calibrate, and callers must fall back to the base rate. There is no code path that surfaces the uncalibrated number.

Shipped

  • [redacted] — calibrated p_at_least_one(), base_rate(), is_established(). Trains strictly on seasons before the one being scored.
  • UnifiedAccuracyTracker.get_td_probability_metrics() + a dashboard section scoring TDs by log loss / skill / AUC instead of a meaningless MAE, with unproven stats visibly labelled.
  • scripts/unit/test_td_probability.py (12 tests).

On 2025 the dashboard now reads: rushing +5.3% skill (AUC 0.673), receiving +5.2% (AUC 0.667), passing −5.1% (AUC 0.606, labelled unproven) — i.e. the flag and the live number agree.

What this does NOT claim

The TD point projections are still bad and this changes nothing about them or about fantasy-point MAE. This is a second, separate output derived from the same λ. Nobody should read "+5.3% log-loss skill over the base rate" as a betting edge — it has not been compared to a market price, and TD props carry heavy vigvigThe bookmaker's built-in cut. It is why a coin-flip bet at -110 needs you to win about 52.4% of the time just to break even..