Compared to the alternatives
Measured against 538, Cook, Sabato, RealClearPolitics and Inside Elections — including every way this is worse, and what it would take to close each gap.
How this compares to the published forecasters, and what it would take to close the gap
Date: 2026-09-29 · single_race_latent/0.7.0 (race model) · house_prior/0.9.0 (House, never promoted)
Written because the honest answer to "is this as good as Silver's work" is no, and not close on the things that matter most. The reasons are specific enough to be a work programme rather than a verdict.
Backtest numbers come from experiments/race_raters.json, experiments/race_raters_governor.json,
experiments/fivethirtyeight_head_to_head.json and experiments/calibration_decomposition.md. Live
2026 numbers for other forecasters were read from their public pages on 2026-09-29 and are cited
where they appear; several are paywalled and are marked as such rather than guessed.
1. The two that decide it
1.1 No live forecast has ever been resolved
Every number in this document is a walk-forward backtest: the model is refitted using only data available before each cycle and scored on that cycle, over ten cycles and 435 statewide races, against 538's own 2018 per-race numbers and fourteen other raters on the same races. That is the algorithm being tested on previous elections, and it is the evidence for everything in §2.
What it is not is a record an outsider can verify, and the gap matters for two specific reasons:
| resolved live forecasts | |
|---|---|
| Nate Silver (538 2008–2022, Silver Bulletin since) | every cycle since 2008, scored by others |
| Cook, Sabato, Inside Elections | decades |
| RealClearPolitics | two decades |
| DDHQ, The Economist, Split Ticket, Race to the WH | several cycles each |
| this project | none. 2026 is the first. |
- A reader cannot check that the backtest was not shaped by knowing the outcomes. The holdout ledger and pre-registration exist to prevent that, but they ask to be trusted; a sealed forecast scored in public asks nothing.
- A backtest tests the algorithm on clean historical data, not the pipeline under live conditions. The defect in §1.2 is that kind of failure: the algorithm handles redrawn districts, and the ingestion never told it 2026 had any.
1.2 The live House number is the outlier, and the reason is a modelling choice
| 2026 House, as of 2026-09-30 | P(D majority) | D seats |
|---|---|---|
this site (house_prior/0.9.0, research) |
99% | 259 [236, 284] |
| this site, at its own measured lean of +2.1 pts | 97% | 250 |
| Split Ticket / The Argument | 90% | — |
| DDHQ | 62% at launch (July 31); 76% reported by The Hill on 2026-09-29 | — |
| Silver Bulletin | paywalled | about a 17-seat Democratic gain per Brookings's summary, so roughly 232 |
| Cook | ratings only | — |
On 2026-09-29 this number was withheld because it had been computed on the 2024 district map (FM-131). That is fixed: all 176 redrawn districts now carry a 2024 presidential baseline on their new lines — Texas and North Carolina from their legislatures' reports, Utah and California from block-level votes, and Florida, Ohio, Tennessee, Louisiana and Alabama by county-to-block areal allocation with a per-district error measured where the truth is known (sd 3.7 points for a district that follows county lines, 13 for one carved from a split county). Incumbency, which had been recorded for no 2026 candidate, is now set for 345 from the dossiers.
The corrected map moved the total from 253 to 259. The map was never what made it high. What makes it high is one assumption: a uniform swing from the realised 2024 House vote (R+2.7) to the generic ballot today (D+9.6), applied to every district. That is +12.3 points everywhere, and it carries the polls' 2024 miss (they read D+0.1 against R−2.6) straight through. The backtest supports uniform swing on past cycles; the other forecasters temper it with district polls, elasticity, candidate quality and ratings, all of which this model lacks. So the number is shown as the model's, badged research, with the sensitivity beside it; the way to a different number is a better model through the gate, not a hand on the dial.
What the model does not use, and the district dossiers (§2.11) now show it should: 173 district polls (FM-132), eight incumbents beaten in primaries, 81 open seats, and unanimous rater calls against it in TX-9, TX-32, TX-35, NC-6, TN-9 and WI-8.
2. Against each forecaster
Two notes on reading the rating rows. Senate covers 2018–2024 and governors the cycles listed.
Every AUC gap now carries a paired bootstrap 95% interval over the shared races
(experiments/race_raters.json, auc_gap_ci95). On these race counts only one gap is
distinguishable from zero: the model beats RealClearPolitics (+0.022, CI +0.006 to +0.042 on
108 Senate races). Every other comparison — Cook, 538, Inside Elections, Sabato, DDHQ, The
Economist, and all governors — is a statistical tie, and the per-forecaster lines below now say
so where yesterday's draft ranked on the third decimal.
2.1 Nate Silver (538 through 2022; Silver Bulletin)
538 itself was closed by ABC in March 2025, so the 538 rows below are its historical record, not a 2026 competitor.
Where this is better
- It is free. Silver's 2026 probabilities and seat counts are behind a paywall.
- It is reproducible to the byte, with run id, commit, seed and archived source bytes for every number. Silver publishes a methodology, not the code or the inputs.
- It publishes a register of 130 of its own defects. Silver does not.
- It pre-registers model changes with a decision rule and keeps a holdout ledger.
- It publishes every fitted parameter in force (
/methodology/parameters). - On 2018 Senate and governor races it scores a better Brier than all three 538 versions (0.0479 against 0.0501–0.0610). The difference is not significant (p = 0.20–0.82 on 59 races), and 2018 is this model's best cycle. Call it a tie.
Where Silver is better
- A resolved public record in every cycle since 2008.
- Senate ranking: 538 AUC 0.9884 against 0.9860 — a tie (gap −0.003, CI −0.012 to +0.006).
- Governors, on the same 37 races: 538 0.9897 against 0.9706 (gap −0.019, CI −0.054 to 0.000): at the edge of distinguishable, and the largest gap against any rater.
- Pollster ratings as a separately validated product; this project has house effects only.
- Candidate quality, fundraising, incumbency detail and expert ratings as model inputs ("deluxe").
- District-level modelling of unpolled House seats from similar districts. This project gives every district the same national swing.
- The 2026 district maps. This project does not have them (§1.2).
- Poll intake: essentially every public poll, against a national shift here that rests on 12 generic-ballot polls from 7 firms, and zero House district polls.
- Written analysis that explains why a number moved.
- Eighteen years of public scrutiny by statisticians.
2.2 Cook Political Report
Where this is better
- Probabilities where Cook has ordinal ratings.
- It calls every race. Cook declined 21 of 107 Senate races as tossups, and this model called 13 of those 21 correctly.
- Senate AUC 0.9860 against 0.9811 and governors 0.9828 against 0.9671 — both ties on the bootstrap (CIs −0.006 to +0.018 and −0.005 to +0.045).
- Every call is traceable to data and a fixed rule.
Where Cook is better
- Zero wrong calls on the races it did call: 0 of 86 Senate and 0 of 38 governor, against this model's 8 and 4.
- Candidate interviews, recruitment, retirements, fundraising and local reporting, none of which exist here.
- The Cook PVI and 2026 district maps.
- Every House race rated individually by people who know the district.
- Decades of record.
2.3 Sabato's Crystal Ball (UVA Center for Politics)
Where this is better
- Probabilities rather than ratings.
- Reproducibility and a published defect log.
- Otherwise nothing measurable: governor AUC is identical (0.9828), and Senate is 0.9860 against 0.9865.
Where Sabato is better
- It calls nearly every race (1 Senate tossup in four cycles) with 5 wrong out of 106, against this model's 8 of 107. That is the same aggressiveness with fewer misses.
- Governors: 2 wrong of 49 against 4 of 51.
- The redistricting analysis this project lacks.
- Qualitative district knowledge and a long record.
2.4 RealClearPolitics
Where this is better
- Better at ranking races, and this is the one gap the bootstrap confirms: Senate 0.9860 against 0.9643 (+0.022, CI +0.006 to +0.042); governors 0.9828 against 0.9616.
- It adjusts for house effects. RCP's average is an unweighted mean of recent polls.
- It gives uncertainty. RCP gives a point average and a map.
- It declines no races. RCP declined 29 Senate races, and this model called 21 of those correctly.
Where RCP is better
- Poll coverage and speed: RCP lists public polls within hours, and this project has thin race polling.
- Zero wrong calls on the 79 Senate races it did call.
- Breadth: approval, the generic ballot and every competitive race in one place.
2.5 Inside Elections
Where this is better
- Probabilities, reproducibility and the defect log.
Where Inside Elections is better
- The best Senate AUC in the table (0.9888) and 4 wrong of 100 called; governors 0.9882 against 0.9828 with 1 wrong of 46. Both gaps are ties on the bootstrap.
- Candidate-level reporting and 2026 maps.
2.6 Decision Desk HQ
Where this is better
- Openness: DDHQ publishes neither code nor parameters.
- Its launch estimate of 62% for the House was far below this model's 99%. With 2026 unresolved, that gap is not evidence for this model.
Where DDHQ is better
- Senate ranking on the same 53 races (2020, 2022): 0.9884 against 0.9812, with 2 wrong against this model's 5 (gap −0.007, CI −0.024 to +0.003): not distinguishable on 53 races.
- Its inputs include candidate quality, fundraising and redistricting.
- It runs a vote-count and race-calling operation, so its result data is first-hand.
2.7 The Economist
Where this is better
- The pre-registration discipline, the holdout ledger and the defect log.
Where The Economist is better
- Senate AUC 0.9841 against 0.9812 on the same 53 races, with 3 wrong against 5.
- Published model code for its 2020 presidential forecast, so its openness is a partial match for this project's.
- Fully Bayesian, with an explicit prior on fundamentals, done by a team.
2.8 Split Ticket (now at The Argument)
Not in the ratings history, so not measured here.
Where Split Ticket is better
- It commissions its own polling. No aggregator can match that, and it is the direct answer to the gap in §3.2.
- District-level modelling on the 2026 maps.
- 90% for the House, which is in line with the other forecasters. This model's 99% is not.
2.9 Race to the WH
Not in the ratings history. Its inputs include polls, historical trends, candidate quality and fundraising, with 50,000 simulations a day. It self-reports second-best accuracy since 2022; that has not been checked here.
Where Race to the WH is better: candidate quality, fundraising and the 2026 maps.
2.10 The smaller raters
On the races they share, this model's AUC is equal to or better than Politico's, Daily Kos's, CNN's and the New York Times's, and better than CBS's, CNalysis's and Election Daily's. It is worse than Fox News on governors (0.9755 against 0.9816). Every one of these samples is between 25 and 82 races, too small to rank.
2.11 What no comparator offers
- Pre-registration with a gate: about 20 challengers evaluated, 2 promoted, the rest refused with their numbers recorded.
- A holdout ledger that marks a cycle burned after 20 specifications have been evaluated on it.
- A register of 130 defects, each with its cost and the thing that now prevents it.
- Byte-level reproducibility from archived sources.
- Refusals by name: 28 House districts withheld with a reason each, and Senate control given as a range where three independents have not declared a caucus.
- Every parameter in force published.
- No LLM ever emits a number.
- A measured profile for every district (
/districts/<race>, "The district in numbers"): eighteen ACS measures on the lines in force, beside the national figure, each a named column of a named Bureau table archived before it was read. Whether the model should use them is a registered question (experiments/preregistration_profile_residuals.md). - A dossier for every one of the 435 districts (
/districts/<race>): the district's section of its state's page read by a language model, each extracted fact stored beside the verbatim sentence it came from, every quote verified against the archived bytes and every number in the brief verified against the inputs before storage. Cook and Sabato do this reading in their heads; nobody publishes it with the citations.
These make the project checkable. They do not make it right more often, and §1.2 shows a checkable project can still publish a wrong headline.
3. Every way this is worse, consolidated
- No resolved live record. Only November fixes this.
- The House headline uses the wrong district map. There is no 2025–26 redistricting (§1.2).
- The House model is unpromoted yet sets the site's most prominent number.
- The House swing is extrapolated past the backtested range, with uniform swing and no elasticity.
- Data operations. The national shift rests on 12 polls from 7 firms, with zero House district polls, and 56 of 506 2026 races with enough polling to forecast.
- No district intelligence: no candidate quality (the derivable part was tested and refused, FM-127), no fundraising, no elasticity, no demographics or turnout.
- Not ahead of the sharpest raters on ranking: every gap against Inside Elections, 538, Sabato, DDHQ and The Economist is a tie on the bootstrap; the only confirmed gap is against RealClearPolitics, in this model's favour.
- More outright misses than any rater that calls as many races (8 of 107 Senate against Sabato's 5 of 106).
- Narrow coverage: no primaries, specials, ballot measures, state legislatures or runoffs.
- No pollster ratings product.
- No analysis, only a deterministic bulletin.
- Ten scored cycles. Cycle-level claims are noisier than their decimals, and the power analysis detects r² = 0.20 only 54% of the time.
- Uncertified House history. House results come from Wikipedia state pages, because MEDSL's
file sits behind a guestbook (
house.py). - One reviewer, and it is an LLM. No hostile human has read this.
4. Action plan
Done on 2026-09-29/30: the 2026 district map (item 1), the unpromoted number badged and, while wrong, withheld (2), district polls ingested (part of 5), dossiers for all 435 districts, and incumbency set from them. What remains, ordered by what changes the House number most:
1. Use the district polls — tried and refused, three ways. house_prior/1.0.0 (polled
districts take their fitted posterior), 1.1.0 (a prior-centred blend at any weight) and
1.2.0 (non-partisan polls only) all lose to the prior-only chamber on the 2018–2024 seat total
(MAE 8.3, ≥6.7, 8.1 against 6.0), with the same signature every time: 2018 pulled Republican,
2020 Democratic — the cycle's national polling miss, which the chamber already prices once as a
shared error. experiments/house_polled_posteriors.md, house_polled_blend.md,
house_polled_nonpartisan.md. The polls stay on the district pages.
2. Historical baselines for the mid-decade redraws, by the areal method with the earlier census blocks, so the backtest exercises the path 2026 takes. Until then the four-cycle MAE (6.0 seats) under-tests it.
3. Open-seat adjustments — tried and refused twice. house_prior/1.5.0 (fixed incumbency
coefficient) and 1.6.0 (walk-forward coefficient, ≈8 points) were each better in 2018, 2022
and 2024 and worse overall (MAE 6.4 and 6.9 against 6.0), because 2020's open seats moved the
wrong way in a year whose error was already Democratic. experiments/house_open_seat*.md.
4. District elasticity — tried and refused. house_prior/1.3.0, each district's fitted
share of the national swing (mean 0.78), is worse than uniform swing in three of four cycles
(MAE 7.7 against 6.0), most of all in 2018, where damping the swing under-predicts the wave
further. experiments/house_elasticity.md. The uniform +12.3 stays because nothing registered
has beaten it, not because it is right.
5. Widen for the extrapolation past the validated swing range, pre-registered.
6. Freeze and pre-register November's scoring — done. First freeze 2026-10-01 (sha256
14308df3…), daily freezes and daily archives of every comparator's published numbers from the
scheduler; the last freeze before election day is scored by the rules in
experiments/preregistration_november_scoring.md.
House first-champion test — refused. Given the same national environment, the
walk-forward seats-votes law (one slope, per-map intercept) has seat MAE 5.83 on 2018–2024; the
district model 6.02. The House card keeps its research badge, and the seats-votes check printed
beside it is an equal, not a sanity check. experiments/house_first_champion.md.
7. Confidence intervals on the rater AUC table — done. Paired bootstrap over shared races;
§2 and /accuracy carry them.
8. Special elections as evaluation folds; 9. pollster ratings as a scored product; 10. a hostile human reviewer.
5. The summary
On ranking statewide races this is competitive: it beats RealClearPolitics by a margin the paired bootstrap confirms and is statistically tied with Cook, Sabato, Inside Elections, 538, DDHQ and The Economist on the races they share, and cannot be separated from 538 on the one cycle with a head-to-head. On breadth, data operations, district intelligence and track record it is far behind all of them.
Its live House number is currently wrong for a known reason: it is forecasting 2026 on 2024 district lines. That is the first thing to fix, and it should be fixed before the forecast is frozen for November.
What it is genuinely better at is being checkable. The defect in §1.2 was found by reading the project's own code, and the fact that it could be found that way is the point of the project.