Election LabU.S. forecasts, with every number traced to its run

Accuracy

walk-forward record of single_race_latent/0.8.1 over 230 scored races

Every figure on this page is walk-forward: each cycle was forecast using only what was estimable before it, and scored against the certified result. Nothing here is fitted to the races it reports on. Champion specification single_race_latent/0.8.1, cycles 2016,2018,2020,2022,2024.

The record

ForecastnMAELean50%80%90%
This model2304.56+1.6447.4%83.9%94.3%
A recency- and size-weighted polling average, same races2305.51+2.86———

The model is 0.95 points better than the polls it is built from, and leans 1.64 points toward the Democrats where the polling average leans 2.86 — so it removes part of the polling error rather than adding to it. That is the comparison that matters: a model that cannot beat its own inputs is an expensive way to repeat them.

Brier score 0.0543, log loss 0.1710 on the same 230 races.

Against the people who do this for a living

Two numbers, because they mean different things. Wrong calls counts only races a rater actually called — a rater who rates more races tossup has fewer chances to be wrong, so the column is not comparable across raters on its own. AUC uses the full ordinal ranking including the tossups. The model calls every race; it has no tossup category.

Senate races, single_race_latent/0.7.0

RaterCyclesRater wrongRated tossupModel right on thoseRater AUCModel AUCVerdict · model − rater [95% CI]
RealClearPolitics2018–20240 / 792921 / 290.96430.9863model · +0.022 [+0.006, +0.042]
Cook Political Report2018–20240 / 862113 / 210.98110.9860tie · +0.005 [-0.006, +0.018]
5382018–20245 / 10073 / 70.98840.9860tie · -0.003 [-0.012, +0.005]
Inside Elections2018–20244 / 10074 / 70.98880.9860tie · -0.003 [-0.012, +0.005]
Sabato's Crystal Ball2018–20245 / 10611 / 10.98650.9860tie · -0.001 [-0.012, +0.010]
Politico2018, 2020, 20220 / 67158 / 150.98330.9833tie · +0.000 [-0.011, +0.011]
Fox News2018, 2022, 20240 / 69127 / 120.98910.9894tie · +0.000 [-0.009, +0.009]
Daily Kos2018, 20200 / 45105 / 100.98400.9867tie · +0.003 [-0.012, +0.020]
Decision Desk HQ2020, 20222 / 5030 / 30.98840.9812tie · -0.007 [-0.024, +0.003]
The Economist2020, 20223 / 5120 / 20.98410.9812tie · -0.003 [-0.018, +0.011]
CNN20180 / 2275 / 70.96670.9778tie · +0.011 [-0.024, +0.048]
New York Times20180 / 19108 / 100.93060.9778tie · +0.047 [-0.005, +0.116]
CBS News20220 / 2342 / 40.99181.0000tie · +0.008 [+0.000, +0.034]
CNalysis20240 / 2321 / 21.00001.0000tie · +0.000 [+0.000, +0.000]
Election Daily20242 / 2500 / 00.99361.0000tie · +0.006 [+0.000, +0.029]

Ratings are read from Wikipedia's transcription of each rater's published table, not from the rater. Where the two disagree this is the one that is wrong. Concurrent same-state contests and special elections are dropped rather than joined by state name alone. A rating is ordinal and is never converted to a probability, so no Brier or log loss is reported against a rater.

Governor races, single_race_latent/0.7.0

RaterCyclesRater wrongRated tossupModel right on thoseRater AUCModel AUCVerdict · model − rater [95% CI]
Cook Political Report2018–20240 / 38139 / 130.96710.9828tie · +0.016 [-0.005, +0.045]
Inside Elections2018–20241 / 4652 / 50.98820.9828tie · -0.005 [-0.026, +0.009]
RealClearPolitics2018–20240 / 371410 / 140.96160.9828tie · +0.021 [-0.005, +0.056]
Sabato's Crystal Ball2018–20242 / 4922 / 20.98280.9828tie · +0.000 [-0.015, +0.017]
Politico2018, 2020, 20220 / 3695 / 90.97980.9798tie · +0.000 [-0.023, +0.023]
Fox News2018, 20220 / 3384 / 80.98160.9755tie · -0.006 [-0.042, +0.020]
5382018, 20221 / 3431 / 30.98970.9706tie · -0.019 [-0.054, +0.000]
Daily Kos2018, 20200 / 2664 / 60.98180.9879tie · +0.006 [-0.018, +0.035]

Ratings are read from Wikipedia's transcription of each rater's published table, not from the rater. Where the two disagree this is the one that is wrong. Concurrent same-state contests and special elections are dropped rather than joined by state name alone. A rating is ordinal and is never converted to a probability, so no Brier or log loss is reported against a rater.

Head to head with 538, 2018

Versionn538 BrierModel Brier538 wrongModel wrongp
538 lite590.06100.0479540.20
538 classic590.05660.0479540.38
538 deluxe590.05010.0479340.82

Read the p column before the Brier column. The model scores better against all three 538 versions and none of the differences is statistically distinguishable from zero. On one cycle of 59 races this comparison cannot separate them, and reporting the Brier difference without that would be claiming a win the data does not support.

2018 is the only cycle 538 published a per-race review for, and it is this model's best cycle. Cook and RealClearPolitics are prohibited in source_registry; Sabato and Silver Bulletin have no registered licence-clean bulk source. 538 'deluxe' includes Cook's and Sabato's ratings as inputs, which is the nearest available comparison to them.

Where the error is

By office

By officenMAELean50%80%90%
governor844.63+0.6746.4%84.5%96.4% ✗
senate1464.52+2.2147.9%83.6%93.2%

By region

By regionnMAELean50%80%90%
MW604.42+2.9045.0%90.0% ✗95.0%
NE404.44-0.2155.0%82.5%95.0%
S694.96+3.2247.8%82.6%94.2%
W614.32-0.1544.3% ✗80.3%93.4%

The regional spread is mostly polling density and the national lean, not geography: the fitted regional term is 0.00 and cannot detect a regional sd below about 2 points.

By how much the race was polled

By how much the race was pollednMAELean50%80%90%
1-4386.33+2.9652.6%89.5% ✗94.7%
10-30884.42+0.9742.0% ✗77.3%93.2%
30-100000513.45+1.4049.0%96.1% ✗100.0% ✗
4-10534.59+2.0750.9%79.2%90.6%

The sparse column is where the model is weakest, and it is why the blend was refitted in 0.7.0 — safe-state polls are wrong by far more than their sample sizes imply.

By how close the race was

By how close the race wasnMAELean50%80%90%
0-3344.02-0.3541.2% ✗88.2% ✗97.1% ✗
20+735.84+3.3146.6%76.7%87.7%
3-8434.01+0.9646.5%83.7%97.7% ✗
8-20803.92+1.3551.2%88.8% ✗97.5% ✗

Error and lean are both largest where the race is least close. The obvious reading — that the forecast is compressed toward the middle — is wrong and was tested: compression would make the lean negative where the Democrat is ahead, and it is positive in seven of eight cells. What this column is showing is the partisan gradient below, seen through a column that happens to correlate with it (experiments/lopsided_or_partisan.md).

By cycle

By cyclenMAELean50%80%90%
2016435.76+4.6825.6% ✗69.8% ✗86.0%
2018633.52-0.5258.7% ✗90.5% ✗100.0% ✗
2020356.25+5.4431.4% ✗71.4% ✗88.6%
2022544.87-0.6046.3%85.2% ✗94.4%
2024352.78+1.4871.4% ✗100.0% ✗100.0% ✗

What is still wrong with it

This section is not an appendix. A track record that reports only the wins is advertising, and every item here is measured, reproducible and currently uncorrected.

The full list runs to 132 entries in docs/FAILURE_MODES.md, each one a defect found in this system, what it cost, and what stopped it recurring.