The method

Where the math takes us.

Mathematical and statistical analysis of collected data — every source cited. Three choices that make our forecast different; the diagrams are illustrative, the commitments are real.

Method 01

Poll misses move together — the simulation moves with them

Schematic · illustrative

the polls saidnational miss

Independent errors

Races miss on their own. The seat distribution comes out absurdly tight — and confidently wrong in years like 2016 and 2020.

Correlated errors

A national miss shifts every race together. We fit the amounts from 20,466 past polls — a national miss of about 3.3–3.8 points, and a race-specific part of about 4.0–4.6 — fitted separately for each chamber.

What  Every forecast is 40,000 simulated elections. When polls miss, they miss together — in 2016 the state errors ran about five points in one direction, in 2020 about four — so the simulation moves races together instead of pretending each race misses alone. The two amounts come from the data, not our judgment: a national miss of about 3.3–3.8 points, and a race-specific part of about 4.0–4.6, fit separately for each chamber from 20,466 historical polls in the FiveThirtyEight archive.

Models that treat every race's error as independent produce seat counts that are far too tight — and get confidently wrong in years like 2016 and 2020, when the polls missed in one direction.

Why it matters

Against the field · 2 comparisons

FiveThirtyEight (classic)

whyalwaysrose / 538 methodology review

They — Also simulated correlated errors — credit where due. But the correlation scales were largely asserted judgment calls.

We — We fit the national vs. race-specific split per chamber from the archive. Independent analysts found asserted scales ran about 14% too narrow; fitting closes that gap.

The Economist

The Economist election model (open source); Linzer 2013 dynamics; Gelman martingale critique

They — Fully Bayesian and open-source (built in R and Stan): a dynamic fundamentals-plus-polls model with state-correlated errors. The principled version of this idea.

We — Ours is the simulation-based analogue — the same honest uncertainty, but with per-chamber scales fitted from history, continuous refit as polls land, and outputs you can audit race by race without statistics software.

Method 02

Candidate quality, WAR-style — on all 506 races

Schematic · illustrative

seat predictsquality scoreactual − expectedexpected margin — fundamentals (pts)actual margin (pts)

A fundamentals model predicts each race's expected result; the gap — the residual — is the candidate's quality score. Applied uniformly to all 506 races — not just the marquee ones. Sabato's 2022 lesson: "choice" beat "referendum."

What  A fundamentals model — how the seat leans, who's incumbent, the demographics, the money — predicts each race's expected margin. The leftover is the candidate's quality score: how much they over- or underperform what the seat alone would predict. It enters our combined forecast as one input per race — on all 435 House races, ~35 Senate seats, and 36 governors' races, not just the marquee ones.

Sabato's 2022 post-mortem is the motivating lesson: 'choice' beat 'referendum' — weak and extreme nominees underperformed the national environment, and top-down models that ignore who is actually running get surprised.

Why it matters

Against the field · 3 comparisons

Split Ticket

Split Ticket WAR models (2024)

They — Lakshya Jain's WAR (wins above replacement) models for the 2024 House and Senate are the most developed public version of residual candidate scoring.

We — Credit to Split Ticket for pioneering it publicly. We implement the same residual idea inside our full ensemble — backtested across 2018 and 2022, applied uniformly to every race.

FiveThirtyEight Deluxe

538 Deluxe methodology

They — Included candidate experience and fundraising as fundamentals inputs — checkbox-style adjustments.

We — We go a level deeper: a continuous score per candidate, estimated from how much they actually over- or underperformed, rather than sorting candidates into experience buckets.

Cook / Inside Elections / Sabato

Sabato's Crystal Ball 2022 post-mortem

They — Assess candidate quality through analyst judgment — and they are very good at it.

We — Ours is systematic and quantitative: the same ruler on every race, no marquee-race bias, and the score is backtested rather than vibes-based.

Method 03

Blending polls and fundamentals — tested, then frozen

Schedule · illustrative shape

Frozen · pre-registered
capped — polls neverfully replace fundamentalsfundamentalspolling12mo out9mo out6mo out3mo outElection Day0%50%100%

The ramp was chosen by testing alternatives against the 2018 and 2022 elections — the best-calibrated schedule won, and it was frozen before the cycle. No October re-tuning.

What  Polls and fundamentals are blended with an explicit, published weight schedule: fundamentals dominate a year out, polling weight ramps to a capped share by Election Day (polls never fully replace fundamentals — late surprises are real). The schedule itself was chosen by testing alternatives against 2018 and 2022 — we tried the options, measured which schedule calibrated best, froze it, and pre-registered it before the cycle.

Everyone blends polls and fundamentals; the blend schedule is where overfitting hides.

Why it matters

Against the field · 2 comparisons

FiveThirtyEight (classic)

538 methodology; our 2018/2022 ablation backtests (pending publication)

They — The same skeleton: polls gain weight toward Election Day, fundamentals never fully zeroed for congressional races.

We — Our schedule was chosen by testing alternatives against 2018/2022 and frozen before the cycle, not tuned by judgment mid-cycle. Same skeleton, with a schedule you can audit.

The Economist

The Economist model; Linzer (2013)

They — Bayesian partial pooling via a dynamic time-series model (Linzer 2013) — the weighting emerges elegantly from the model structure.

We — Ours is the combined-forecast analogue: explicit, published weights instead of emergent ones. You can read our schedule in one table — no statistical sampler required.

FiveThirtyEight built the modern forecasting stack and published enough of it to learn from. The Economist proved a fully open Bayesian model could compete with the best closed ones. Independent analysts — whyalwaysrose, grantbw4, and the Race to the White House team — found the failure modes everyone else missed. Second Opinion stands on all of it; the methods above are ours, the foundation is theirs.