New to these reports? Start here
  • Dotted-underlined words have a plain-English definition — hover or tap them. Every term is also on the glossary page.
  • "Null" means we found nothing, not that something broke. Most reports here are negative results, on purpose — knowing an idea doesn't work is the point.
  • Two questions get asked separately. First, is the effect real? Second, is it already priced into the betting odds? An effect can be completely real and still useless to bet on.
  • A "calibration" row is a self-check. It runs the same method on something already known to be true. If that fails, the whole report is unreliable — so it's reported alongside the findings.
  • If a confidence interval includes zero, the real effect might be nothing at all, so no claim gets made.

How well does the public pick? — field baseline, 2014–2025

2026-09-13. field_baseline.py → field_baseline_results.json (gitignored). Source: NFL Pickwatch per-game pick shares, captured by scripts/nfl/capture_pickwatch.py into [redacted] (weekly cron, Tue 14:00 UTC). Lines and results come from nflverse, not Pickwatch, whose spreads are sometimes wrong (2025 W1 TB@ATL: ATL +7 next to a +111 moneyline).

3,142 regular-season games, 12 seasons, 100% matched. Per game, a median of ~5,000 fan picks and ~300 media-expert picks. Ties excluded (no side is right).

⚠ Pickwatch's terms allow personal, non-commercial use only and forbid republishing. These numbers stay in backtests/; nothing here goes on a public page.

Sources considered

Source Status
NFL Pickwatch ✅ used. robots.txt permits; JSON API has no robots file
Yahoo Pick Distribution ❌ terms forbid automated collection without permission; robots.txt disallows ClaudeBot/anthropic-ai. Our own pool's Yahoo PDF exports stay the manual route
Fantasy Nerds API ❌ free key returns sample data only
PoolGenius ❌ paid

Headline

(pooled, 2014–2025) Straight-up % of max confidence pts*
Always-favorite 66.3% 71.9%
Fan majority vote 66.1% —
Expert majority vote 65.9% —
Average fan pick 63.3% 68.8%
Average expert pick 63.3% 68.9%

* Pickwatch publishes no confidence ranks; the field columns assume the field ranks by spread like the baseline does. See the real-pool check — that assumption flatters the field.

Favorite minus average fan: +3.0pp SU (95% CI95% confidence intervalThe range the true value is plausibly in. If this range includes zero, we cannot rule out that the real effect is nothing at all. +2.1 to +4.0, week bootstrapbootstrapRe-running a calculation on thousands of resampled versions of the data to see how much the answer wobbles. The spread of those answers becomes the confidence interval.). The interval for experts is the same: media experts pick no better than fans.

  • The majority vote is the favorite. Majority accuracy is within 0.2pp of always-favorite. Individuals lose ground by deviating from their own crowd.
  • Not every year. 2021: favorite 62.4%, fans 62.3%. 2017 and 2023: favorite ahead by 6–7pp. The pooled gap is stable; single seasons are not.

The crowd over-concentrates on favorites

Market P(favorite) games favorite won fans on favorite experts on favorite
0.5–0.6 993 56.1% 62.9% 61.9%
0.6–0.7 1,054 62.5% 83.4% 83.0%
0.7–0.8 795 77.0% 93.9% 94.5%
0.8–1.0 289 86.5% 98.1% 98.5%

A 65% favorite draws 83% of picks. For straight-up points that's harmless: picking the favorite is still right. For winning a pool it's the only place differentiation is cheap. But FINDINGS_POOL_2026.md found contrarianism loses in every human-rival regime, so this describes the field; it doesn't license a strategy change.

Real-pool check (our Yahoo pool, 2025 weeks 1–12, 177 games)

value
Pool members on the favorite 81.5%
Pickwatch fans on the favorite, same games 84.4%
Per-game correlationcorrelationHow closely two things move together, from -1 (opposite) through 0 (unrelated) to +1 (in lockstep). It does not by itself mean one causes the other. 0.92 (mean abs diff 5.9pp)
Members' confidence vs spread (median SpearmanSpearman correlationA measure of whether two rankings agree, from -1 (opposite) through 0 (unrelated) to +1 (identical). Cares about order, not exact values.) 0.57
Pool's actual % of max 69.1%
Pool's own picks re-ranked by spread 71.5%
  • Pickwatch fans are a good stand-in for our pool's picks. They lean slightly more to favorites.
  • The spread-ranking proxy overstates a real field by ~2.4pp, because real people don't rank by spread. Applying that correction, a realistic field sits around 66–67% of max, ~5pp behind always-favorite. That matches our pool directly: baseline 74.5% vs field mean 69.3% (FINDINGS_REAL_POOL_2026.md).

Caveats

  • Timing flatters the favorite. nflverse's closing lineclosing lineThe final odds right before a game starts. It reflects everything the betting market knows, which makes it the hardest benchmark to beat. knows everything up to kickoff; many people pick earlier in the week. Some of the 3pp is late information, not skill. In a pool with a Thursday lock the realistic gap is somewhat smaller.
  • Pickwatch fans self-select — they're engaged enough to log picks on a tracking site. The pool check says they resemble our pool, but that is one pool and 12 weeks.
  • Straight-up only. No confidence, no ATS pools (Pickwatch has ATS shares in raw_json; not analysed here).

Next

  • archive_vs_field in the script lines up our archived 2026 sheets (pickem_archive.db, baseline and model) against the field on the same games as weeks get graded. Wait until the season ends before comparing.