New to these reports? Start here
  • Dotted-underlined words have a plain-English definition — hover or tap them. Every term is also on the glossary page.
  • "Null" means we found nothing, not that something broke. Most reports here are negative results, on purpose — knowing an idea doesn't work is the point.
  • Two questions get asked separately. First, is the effect real? Second, is it already priced into the betting odds? An effect can be completely real and still useless to bet on.
  • A "calibration" row is a self-check. It runs the same method on something already known to be true. If that fails, the whole report is unreliable — so it's reported alongside the findings.
  • If a confidence interval includes zero, the real effect might be nothing at all, so no claim gets made.

College hockey model vs the betting market (first comparison, 2026-09-29)

The market is better than the model, and the model adds nothing to it.

Data

  • Odds: OddsPapi historical snapshots, January–April 2026 (the archive starts in January). Source: scripts/college_hockey/oddspapi_backfill.py; historical calls are free.
  • Coverage: 173 of 561 matched D-I games have a moneyline (DraftKings 155, BetRivers 56, Pinnacle 51, Caesars 18).
  • The "close" is not a true close. It is the last stored snapshot before puck drop, preferring Pinnacle, then DraftKings. The median is 6 hours before the start and the median overround is 1.062.
  • Models: the clean walk-forwardwalk-forwardEvaluating week by week using only what was knowable before each week, mimicking how the model would actually have been used at the time. fold (weights trained on seasons < 2025, EloElo ratingA rating system, originally from chess, that moves a team up or down based on results and the strength of the opponent. as of each game date). The deployed model is scored only from 2026-02-14, after its weights were trained.

Results (ties dropped; model_vs_market_results.json)

n log losslog lossA score for probability forecasts that punishes confident wrong answers harshly. Lower is better. model log loss market gap (95% CI95% confidence intervalThe range the true value is plausibly in. If this range includes zero, we cannot rule out that the real effect is nothing at all.) accuracy model / market
Walk-forward 164 0.625 0.603 +0.021 (−0.014, +0.057) 64.0% / 68.9%
Deployed, from Feb 14 89 0.634 0.595 +0.040 (−0.009, +0.087) 66.3% / 71.9%
  • Blending: the out-of-fold blend of market and model is worse than the market alone (0.612 vs 0.603 walk-forward).
  • Betting: backing the model where it disagrees by ≥3 pp gives 133 bets at −16.6% ROIreturn on investmentProfit as a percentage of the money wagered. +2% means $2 profit per $100 bet. (CI −38% to +4%).

⚠ The first run said the opposite, and it was wrong

The first run showed the model beating the market (+34% ROI, CI above zero). It was caused by 33 games in the games table with the wrong winner: home/away were swapped while the scores stayed where they were.

  • Nearly all were conference-playoff games. CHN prints a game number before playoff results ("06 Fri 2 W 3 - 0 Miami"). The rescrape pattern didn't allow for it, so those rows were never parsed, and both teams' pages claimed to be home.
  • The tell was the market scoring worse than a coin flip (log loss 0.706).
  • The fix corrected all 33 and completed 21 unscored playoff games. The rescrape now corrects any completed row whose winner disagrees with CHN.
  • 2024-25 audited clean.
  • The site's forward record moved from 58.7% to 60.5% (n=334) once results were right.

On this page

Terms in this report

Related

College hockey betting angles backlog (2026-09-29) Pre-registration: roster turnover vs the October market (2026-27)

Source

backtests/college_hockey_eval/FINDINGS_MODEL_VS_MARKET.md
updated 2026-09-29 15:28