New to these reports? Start here
- Dotted-underlined words have a plain-English definition — hover or tap them. Every term is also on the glossary page.
- "Null" means we found nothing, not that something broke. Most reports here are negative results, on purpose — knowing an idea doesn't work is the point.
- Two questions get asked separately. First, is the effect real? Second, is it already priced into the betting odds? An effect can be completely real and still useless to bet on.
- A "calibration" row is a self-check. It runs the same method on something already known to be true. If that fails, the whole report is unreliable — so it's reported alongside the findings.
- If a confidence interval includes zero, the real effect might be nothing at all, so no claim gets made.
How well does the public pick? — field baseline, 2014–2025
2026-09-13. field_baseline.py → field_baseline_results.json (gitignored).
Source: NFL Pickwatch per-game pick shares, captured by
scripts/nfl/capture_pickwatch.py into [redacted] (weekly cron, Tue
14:00 UTC). Lines and results come from nflverse, not Pickwatch, whose spreads are
sometimes wrong (2025 W1 TB@ATL: ATL +7 next to a +111 moneyline).
3,142 regular-season games, 12 seasons, 100% matched. Per game, a median of ~5,000 fan picks and ~300 media-expert picks. Ties excluded (no side is right).
⚠ Pickwatch's terms allow personal, non-commercial use only and forbid republishing. These numbers stay in
backtests/; nothing here goes on a public page.
Sources considered
| Source | Status |
|---|---|
| NFL Pickwatch | ✅ used. robots.txt permits; JSON API has no robots file |
| Yahoo Pick Distribution | ❌ terms forbid automated collection without permission; robots.txt disallows ClaudeBot/anthropic-ai. Our own pool's Yahoo PDF exports stay the manual route |
| Fantasy Nerds API | ❌ free key returns sample data only |
| PoolGenius | ❌ paid |
Headline
| (pooled, 2014–2025) | Straight-up | % of max confidence pts* |
|---|---|---|
| Always-favorite | 66.3% | 71.9% |
| Fan majority vote | 66.1% | — |
| Expert majority vote | 65.9% | — |
| Average fan pick | 63.3% | 68.8% |
| Average expert pick | 63.3% | 68.9% |
* Pickwatch publishes no confidence ranks; the field columns assume the field ranks by spread like the baseline does. See the real-pool check — that assumption flatters the field.
Favorite minus average fan: +3.0pp SU (95% CI95% confidence intervalThe range the true value is plausibly in. If this range includes zero, we cannot rule out that the real effect is nothing at all. +2.1 to +4.0, week bootstrapbootstrapRe-running a calculation on thousands of resampled versions of the data to see how much the answer wobbles. The spread of those answers becomes the confidence interval.). The interval for experts is the same: media experts pick no better than fans.
- The majority vote is the favorite. Majority accuracy is within 0.2pp of always-favorite. Individuals lose ground by deviating from their own crowd.
- Not every year. 2021: favorite 62.4%, fans 62.3%. 2017 and 2023: favorite ahead by 6–7pp. The pooled gap is stable; single seasons are not.
The crowd over-concentrates on favorites
| Market P(favorite) | games | favorite won | fans on favorite | experts on favorite |
|---|---|---|---|---|
| 0.5–0.6 | 993 | 56.1% | 62.9% | 61.9% |
| 0.6–0.7 | 1,054 | 62.5% | 83.4% | 83.0% |
| 0.7–0.8 | 795 | 77.0% | 93.9% | 94.5% |
| 0.8–1.0 | 289 | 86.5% | 98.1% | 98.5% |
A 65% favorite draws 83% of picks. For straight-up points that's harmless:
picking the favorite is still right. For winning a pool it's the only place
differentiation is cheap. But FINDINGS_POOL_2026.md found contrarianism loses
in every human-rival regime, so this describes the field; it doesn't license a
strategy change.
Real-pool check (our Yahoo pool, 2025 weeks 1–12, 177 games)
| value | |
|---|---|
| Pool members on the favorite | 81.5% |
| Pickwatch fans on the favorite, same games | 84.4% |
| Per-game correlationcorrelationHow closely two things move together, from -1 (opposite) through 0 (unrelated) to +1 (in lockstep). It does not by itself mean one causes the other. | 0.92 (mean abs diff 5.9pp) |
| Members' confidence vs spread (median SpearmanSpearman correlationA measure of whether two rankings agree, from -1 (opposite) through 0 (unrelated) to +1 (identical). Cares about order, not exact values.) | 0.57 |
| Pool's actual % of max | 69.1% |
| Pool's own picks re-ranked by spread | 71.5% |
- Pickwatch fans are a good stand-in for our pool's picks. They lean slightly more to favorites.
- The spread-ranking proxy overstates a real field by ~2.4pp, because real
people don't rank by spread. Applying that correction, a realistic field sits
around 66–67% of max, ~5pp behind always-favorite. That matches our pool
directly: baseline 74.5% vs field mean 69.3% (
FINDINGS_REAL_POOL_2026.md).
Caveats
- Timing flatters the favorite. nflverse's closing lineclosing lineThe final odds right before a game starts. It reflects everything the betting market knows, which makes it the hardest benchmark to beat. knows everything up to kickoff; many people pick earlier in the week. Some of the 3pp is late information, not skill. In a pool with a Thursday lock the realistic gap is somewhat smaller.
- Pickwatch fans self-select — they're engaged enough to log picks on a tracking site. The pool check says they resemble our pool, but that is one pool and 12 weeks.
- Straight-up only. No confidence, no ATS pools (Pickwatch has ATS shares in
raw_json; not analysed here).
Next
archive_vs_fieldin the script lines up our archived 2026 sheets (pickem_archive.db, baseline and model) against the field on the same games as weeks get graded. Wait until the season ends before comparing.