Accuracy
walk-forward record of single_race_latent/0.8.1 over 230 scored races
Every figure on this page is walk-forward: each cycle was forecast using only what was estimable before it, and scored against the certified result. Nothing here is fitted to the races it reports on. Champion specification single_race_latent/0.8.1, cycles 2016,2018,2020,2022,2024.
The record
| Forecast | n | MAE | Lean | 50% | 80% | 90% |
|---|---|---|---|---|---|---|
| This model | 230 | 4.56 | +1.64 | 47.4% | 83.9% | 94.3% |
| A recency- and size-weighted polling average, same races | 230 | 5.51 | +2.86 | — | — | — |
The model is 0.95 points better than the polls it is built from, and leans 1.64 points toward the Democrats where the polling average leans 2.86 — so it removes part of the polling error rather than adding to it. That is the comparison that matters: a model that cannot beat its own inputs is an expensive way to repeat them.
Brier score 0.0543, log loss 0.1710 on the same 230 races.
Against the people who do this for a living
Two numbers, because they mean different things. Wrong calls counts only races a rater actually called — a rater who rates more races tossup has fewer chances to be wrong, so the column is not comparable across raters on its own. AUC uses the full ordinal ranking including the tossups. The model calls every race; it has no tossup category.
Senate races, single_race_latent/0.7.0
| Rater | Cycles | Rater wrong | Rated tossup | Model right on those | Rater AUC | Model AUC | Verdict · model − rater [95% CI] |
|---|---|---|---|---|---|---|---|
| RealClearPolitics | 2018–2024 | 0 / 79 | 29 | 21 / 29 | 0.9643 | 0.9863 | model · +0.022 [+0.006, +0.042] |
| Cook Political Report | 2018–2024 | 0 / 86 | 21 | 13 / 21 | 0.9811 | 0.9860 | tie · +0.005 [-0.006, +0.018] |
| 538 | 2018–2024 | 5 / 100 | 7 | 3 / 7 | 0.9884 | 0.9860 | tie · -0.003 [-0.012, +0.005] |
| Inside Elections | 2018–2024 | 4 / 100 | 7 | 4 / 7 | 0.9888 | 0.9860 | tie · -0.003 [-0.012, +0.005] |
| Sabato's Crystal Ball | 2018–2024 | 5 / 106 | 1 | 1 / 1 | 0.9865 | 0.9860 | tie · -0.001 [-0.012, +0.010] |
| Politico | 2018, 2020, 2022 | 0 / 67 | 15 | 8 / 15 | 0.9833 | 0.9833 | tie · +0.000 [-0.011, +0.011] |
| Fox News | 2018, 2022, 2024 | 0 / 69 | 12 | 7 / 12 | 0.9891 | 0.9894 | tie · +0.000 [-0.009, +0.009] |
| Daily Kos | 2018, 2020 | 0 / 45 | 10 | 5 / 10 | 0.9840 | 0.9867 | tie · +0.003 [-0.012, +0.020] |
| Decision Desk HQ | 2020, 2022 | 2 / 50 | 3 | 0 / 3 | 0.9884 | 0.9812 | tie · -0.007 [-0.024, +0.003] |
| The Economist | 2020, 2022 | 3 / 51 | 2 | 0 / 2 | 0.9841 | 0.9812 | tie · -0.003 [-0.018, +0.011] |
| CNN | 2018 | 0 / 22 | 7 | 5 / 7 | 0.9667 | 0.9778 | tie · +0.011 [-0.024, +0.048] |
| New York Times | 2018 | 0 / 19 | 10 | 8 / 10 | 0.9306 | 0.9778 | tie · +0.047 [-0.005, +0.116] |
| CBS News | 2022 | 0 / 23 | 4 | 2 / 4 | 0.9918 | 1.0000 | tie · +0.008 [+0.000, +0.034] |
| CNalysis | 2024 | 0 / 23 | 2 | 1 / 2 | 1.0000 | 1.0000 | tie · +0.000 [+0.000, +0.000] |
| Election Daily | 2024 | 2 / 25 | 0 | 0 / 0 | 0.9936 | 1.0000 | tie · +0.006 [+0.000, +0.029] |
Ratings are read from Wikipedia's transcription of each rater's published table, not from the rater. Where the two disagree this is the one that is wrong. Concurrent same-state contests and special elections are dropped rather than joined by state name alone. A rating is ordinal and is never converted to a probability, so no Brier or log loss is reported against a rater.
Governor races, single_race_latent/0.7.0
| Rater | Cycles | Rater wrong | Rated tossup | Model right on those | Rater AUC | Model AUC | Verdict · model − rater [95% CI] |
|---|---|---|---|---|---|---|---|
| Cook Political Report | 2018–2024 | 0 / 38 | 13 | 9 / 13 | 0.9671 | 0.9828 | tie · +0.016 [-0.005, +0.045] |
| Inside Elections | 2018–2024 | 1 / 46 | 5 | 2 / 5 | 0.9882 | 0.9828 | tie · -0.005 [-0.026, +0.009] |
| RealClearPolitics | 2018–2024 | 0 / 37 | 14 | 10 / 14 | 0.9616 | 0.9828 | tie · +0.021 [-0.005, +0.056] |
| Sabato's Crystal Ball | 2018–2024 | 2 / 49 | 2 | 2 / 2 | 0.9828 | 0.9828 | tie · +0.000 [-0.015, +0.017] |
| Politico | 2018, 2020, 2022 | 0 / 36 | 9 | 5 / 9 | 0.9798 | 0.9798 | tie · +0.000 [-0.023, +0.023] |
| Fox News | 2018, 2022 | 0 / 33 | 8 | 4 / 8 | 0.9816 | 0.9755 | tie · -0.006 [-0.042, +0.020] |
| 538 | 2018, 2022 | 1 / 34 | 3 | 1 / 3 | 0.9897 | 0.9706 | tie · -0.019 [-0.054, +0.000] |
| Daily Kos | 2018, 2020 | 0 / 26 | 6 | 4 / 6 | 0.9818 | 0.9879 | tie · +0.006 [-0.018, +0.035] |
Ratings are read from Wikipedia's transcription of each rater's published table, not from the rater. Where the two disagree this is the one that is wrong. Concurrent same-state contests and special elections are dropped rather than joined by state name alone. A rating is ordinal and is never converted to a probability, so no Brier or log loss is reported against a rater.
Head to head with 538, 2018
| Version | n | 538 Brier | Model Brier | 538 wrong | Model wrong | p |
|---|---|---|---|---|---|---|
| 538 lite | 59 | 0.0610 | 0.0479 | 5 | 4 | 0.20 |
| 538 classic | 59 | 0.0566 | 0.0479 | 5 | 4 | 0.38 |
| 538 deluxe | 59 | 0.0501 | 0.0479 | 3 | 4 | 0.82 |
Read the p column before the Brier column. The model scores better against all three 538 versions and none of the differences is statistically distinguishable from zero. On one cycle of 59 races this comparison cannot separate them, and reporting the Brier difference without that would be claiming a win the data does not support.
2018 is the only cycle 538 published a per-race review for, and it is this model's best cycle. Cook and RealClearPolitics are prohibited in source_registry; Sabato and Silver Bulletin have no registered licence-clean bulk source. 538 'deluxe' includes Cook's and Sabato's ratings as inputs, which is the nearest available comparison to them.
Where the error is
By office
| By office | n | MAE | Lean | 50% | 80% | 90% |
|---|---|---|---|---|---|---|
| governor | 84 | 4.63 | +0.67 | 46.4% | 84.5% | 96.4% ✗ |
| senate | 146 | 4.52 | +2.21 | 47.9% | 83.6% | 93.2% |
By region
| By region | n | MAE | Lean | 50% | 80% | 90% |
|---|---|---|---|---|---|---|
| MW | 60 | 4.42 | +2.90 | 45.0% | 90.0% ✗ | 95.0% |
| NE | 40 | 4.44 | -0.21 | 55.0% | 82.5% | 95.0% |
| S | 69 | 4.96 | +3.22 | 47.8% | 82.6% | 94.2% |
| W | 61 | 4.32 | -0.15 | 44.3% ✗ | 80.3% | 93.4% |
The regional spread is mostly polling density and the national lean, not geography: the fitted regional term is 0.00 and cannot detect a regional sd below about 2 points.
By how much the race was polled
| By how much the race was polled | n | MAE | Lean | 50% | 80% | 90% |
|---|---|---|---|---|---|---|
| 1-4 | 38 | 6.33 | +2.96 | 52.6% | 89.5% ✗ | 94.7% |
| 10-30 | 88 | 4.42 | +0.97 | 42.0% ✗ | 77.3% | 93.2% |
| 30-100000 | 51 | 3.45 | +1.40 | 49.0% | 96.1% ✗ | 100.0% ✗ |
| 4-10 | 53 | 4.59 | +2.07 | 50.9% | 79.2% | 90.6% |
The sparse column is where the model is weakest, and it is why the blend was refitted in 0.7.0 — safe-state polls are wrong by far more than their sample sizes imply.
By how close the race was
| By how close the race was | n | MAE | Lean | 50% | 80% | 90% |
|---|---|---|---|---|---|---|
| 0-3 | 34 | 4.02 | -0.35 | 41.2% ✗ | 88.2% ✗ | 97.1% ✗ |
| 20+ | 73 | 5.84 | +3.31 | 46.6% | 76.7% | 87.7% |
| 3-8 | 43 | 4.01 | +0.96 | 46.5% | 83.7% | 97.7% ✗ |
| 8-20 | 80 | 3.92 | +1.35 | 51.2% | 88.8% ✗ | 97.5% ✗ |
Error and lean are both largest where the race is least close. The obvious reading — that the forecast is compressed toward the middle — is wrong and was tested: compression would make the lean negative where the Democrat is ahead, and it is positive in seven of eight cells. What this column is showing is the partisan gradient below, seen through a column that happens to correlate with it (experiments/lopsided_or_partisan.md).
By cycle
| By cycle | n | MAE | Lean | 50% | 80% | 90% |
|---|---|---|---|---|---|---|
| 2016 | 43 | 5.76 | +4.68 | 25.6% ✗ | 69.8% ✗ | 86.0% |
| 2018 | 63 | 3.52 | -0.52 | 58.7% ✗ | 90.5% ✗ | 100.0% ✗ |
| 2020 | 35 | 6.25 | +5.44 | 31.4% ✗ | 71.4% ✗ | 88.6% |
| 2022 | 54 | 4.87 | -0.60 | 46.3% | 85.2% ✗ | 94.4% |
| 2024 | 35 | 2.78 | +1.48 | 71.4% ✗ | 100.0% ✗ | 100.0% ✗ |
What is still wrong with it
This section is not an appendix. A track record that reports only the wins is advertising, and every item here is measured, reproducible and currently uncorrected.
- It leans about two points toward the Democrats. +2.08 over 435 scored races. Five pre-registered attempts to correct the level have been refused by the promotion gate, because the cycle mean has a coefficient of variation near 4 and flips sign in half the cycles measured: nothing estimable before a cycle predicts that cycle's lean. The lean is quoted on every chamber page as a sensitivity rather than silently corrected.
- There is a partisan gradient it does not correct. The polls overstate the
Democrat more in redder states, in proportion to how red the state is — negative in 9 of 9 cycles,
coefficient of variation 0.52, independent of the level. In the reddest quarter of states the lean
is +5.9 points whether the model has the Democrat or the Republican ahead — which side is favoured
makes no difference, only how red the state is. Correcting it improves MAE from 5.46 to
5.24 and 90% coverage from 0.887 to 0.906. It was pre-registered under log loss, which it
does not improve, and the gate refused it (q = 0.685). Switching metric after seeing the result
would make the pre-registration worthless, so it stays uncorrected and marked
suggestive_onlyuntil 2026, its declared confirmatory cycle. - 2014 is the worst cycle and is not fully explained. Bias +6.50, 90% coverage 0.723 against a nominal 0.90.
- The regional term is 0.00 and that is a bound, not a measurement. The estimator cannot detect a regional sd below about 2 points, so "no regional structure" means "none large enough for this design to see".
- The 50% band is the weakest. Measured 0.471 against nominal 0.500 — inside tolerance, but it is the band that has failed before and the one to watch.
- The evidence base is ten cycles. Cycle-level effects have ten observations, not 435. Anything this page reports per cycle is noisier than its decimal places suggest.
The full list runs to 132 entries in docs/FAILURE_MODES.md, each one a defect found in this system, what it cost, and what stopped it recurring.