FutureBallotU.S. election forecasts, with every number traced to its source

The statistical model

Every distribution, every prior and every term in the variance budget, with the assumption each one rests on written next to it.

STATISTICAL MODEL

This document specifies the mathematics. It is written for a statistician who has not read the code and who intends to find the mistakes. Every constant that appears below is either estimated from data, read from configuration, or accompanied by a justification. There are no unexplained magic numbers.

Status, 2026-09-27: fitted, and validated where it says so. This line read "design specification, not yet fitted. Nothing here is validated" until today, which was written before anything was fitted and left standing for months afterwards — so the project's primary statistical document asserted the opposite of the truth. The model has been fitted, walk-forward backtested and scored on 435 races over ten cycles, 2006-2024 (single_race_latent/0.7.0: MAE 4.89, lean +2.08, realised coverage 0.471 / 0.798 / 0.903 against a nominal 0.50 / 0.80 / 0.90 -- all three inside the ADR-0018 floor), and two challengers have been promoted through a pre-registered gate. On those same races a recency- and size-weighted polling average scores MAE 5.51 with a lean of +2.86, so the model beats its baseline and removes 0.77 points of the polls' lean rather than adding to it (F-11, experiments/calibration_decomposition.md).

The five-cycle figures this paragraph used to quote (295 races, MAE 4.68, coverage 0.475 / 0.810 / 0.919) were for 0.6.0 before the evidence base was extended to 2006 and before FM-104's fix; they are not comparable to the ten-cycle ones and are not the same population.

What remains genuinely unvalidated is named where it arises, not hedged here. The cross-race posterior predictive check now passes on all six cycles, 2014 at tail p = 0.089 where it failed at 0.037 (experiments/phase5_validation_2026-09-29.json, experiments/phase5_exit_met.md) -- a pass, and the tightest of the six against a next-closest 0.110. sigma_region fits at 0.00, and that is now a bounded claim rather than a shrug: measured on the real skeleton, the estimator finds a true 2.0-point regional term 98% of the time and a 1.0-point one 38% of the time, so a regional term of 2.0 points or more is ruled out and one of 1.0 or less is undetectable here. The forecast's own regional signal, debiased for the noise on each regional mean, is 0.72 points across all races and 0.00 among well-polled ones — below the floor, so the fitted zero and the slice's 3.33-point spread agree. A spread over four groups is about twice a standard deviation and each regional mean carries ±1.2 of noise, which is where the apparent contradiction came from (experiments/region_tension.md). sigma_office fits at 0.00 on the same estimator and its floor also lands at 2.0 points, but more weakly — 82% detection at a true 2.0 against the regional 98% — so it bounds less. Where a parameterisation below is still a competing candidate rather than a decision, it says so in place.


1. Notation

Symbol Meaning
r race (a contest in one geography for one office in one cycle)
c candidate within a race; K_r candidates plus an "undecided/other" category
t time, in days, with t = 0 at Election Day and t < 0 before it
i an individual poll observation
p(i) pollster of poll i; f(p) its pollster family
n_i nominal sample size
y_i reported vector of candidate shares
θ_r(t) latent population support in race r at time t
D_X the information set knowable at calendar date X

Two quantities of interest, always distinct:

They differ by exactly one term: integrated future opinion drift. §11 makes that precise.


2. Scale

Vote shares are compositional; modelling them on the raw simplex with Gaussian noise is wrong at the boundaries and produces impossible intervals in lopsided races.

We work in additive log-ratio (ALR) space:

φ_rc(t) = log( θ_rc(t) / θ_r,ref(t) ),    c = 1 … K_r - 1

Choice of reference category (revised after adversarial review, AR-10). The reference sits in the denominator of every log-ratio, so it must not approach zero. Undecided is therefore a bad choice: it shrinks toward zero as Election Day nears, and the log-ratios would destabilise precisely in the final fortnight when precision matters most.

The reference is the leading major-party candidate, fixed per race at fit time from the fundamentals prior — chosen from the prior, not from the data being fitted, so the parameterisation does not depend on the likelihood. Undecided remains an explicit modelled category (§9); it is simply not the denominator. The centred-log-ratio (CLR) parameterisation, which avoids choosing a reference at all, is retained as the tested alternative.

For the common two-candidate case this reduces to a single logit of the two-party share, and the margin m_r = θ_rA - θ_rB is recovered by the inverse transform inside the simulation, so all uncertainty propagates to the margin without a normal approximation on the margin itself.

Candidate alternative to test: a Dirichlet-multinomial observation model directly on the simplex. Selection is empirical (ADR-0005), not aesthetic.


3. Latent opinion: the state process

For each race, the latent vector evolves as a Gaussian random walk with a national component:

φ_r(t) = α_r  +  β_r · η(t)  +  u_r(t)

σ_η and σ_u are not chosen by hand. They are estimated from historical polling trajectories: the model is fitted on past cycles where the truth is known, and the innovation scales are hyperparameters with weakly-informative priors updated by that history. The posterior for σ_η is then used as a prior for the live cycle.

3.1 Why a state-space model rather than a weighted average

A weighted average answers "what is the centre of recent polls". A state-space model answers "what is public opinion, given that polls are noisy, biased, and irregularly spaced". The second is the question we actually have. Critically, it borrows strength across time and across races through η(t), so a state with two polls is informed by the national trend rather than by its own two polls alone.

This must still be proved to beat a recency-weighted average out of sample (§13). If it does not, we ship the average.


4. Observation model for polls

For poll i fielded over [t_start, t_end] in race r:

ỹ_i  ~  Student-t_ν ( μ_i , Σ_i )

μ_i = φ_r(t̄_i)  +  h_{p(i)}  +  m_{mode(i)}  +  g_{pop(i)}(t̄_i)  +  s_i · ψ

where

4.1 Variance

Σ_i = DEFF_i · Σ_sampling(n_i, θ)  +  τ²_{p(i)}  +  τ²_{family(p(i))}

4.2 Non-independence

Three mechanisms break the "one poll = one independent observation" assumption, and each is modelled explicitly rather than ignored:

  1. Tracking polls / rolling samples. Consecutive releases share respondents. Overlapping releases from the same tracker are collapsed into non-overlapping windows, or given a correlation term ρ_overlap proportional to the fraction of shared field days.
  2. Shared samples across questions. One survey reporting Senate and Governor toplines is one sample; its errors correlate across races. Modelled with a per-sample random effect shared by all races it reports.
  3. Pollster families. Firms sharing panels, vendors, or methodology (f(p)) get a family-level house effect and family-level variance term, so releasing five polls under five brand names does not buy five independent observations.

5. Identification — the part most implementations get wrong

The decomposition μ_i = φ_r(t) + h_p is not identified without a constraint: adding a constant to every h_p and subtracting it from every φ_r leaves the likelihood unchanged.

Polls alone can never tell us the absolute level of opinion. Only election results can. Therefore:

  1. Within a cycle, house effects are identified only relative to each other. We impose a volume-weighted sum-to-zero constraint on h_p within each cycle so the latent state is pinned to "the average pollster".
  2. The industry-wide bias — the amount by which the average pollster misses the result — is a separate parameter b_cycle, estimable only from realised results across past cycles. It is the dominant term in §10's correlated error.
  3. Consequently, a model fitted on polls alone and reporting tight intervals is reporting the uncertainty of "where the polling average is", not "where the election lands". Conflating these is the single most common failure in public election models. The variance budget in §11 keeps them apart.

Number of effective observations of industry-wide bias is small — roughly one per office-type per cycle, so on the order of 10–20 for the modern polling era. b_cycle's scale parameter is therefore weakly identified and must carry a wide, honestly-reported posterior. This is a fundamental limit of the problem, not of the implementation, and is recorded as failure mode FM-14.


6. Pollster quality, hierarchically

No subjective letter grades. For each pollster p:

h_p      ~ N( μ_h[ mode_p , partisan_p ] , σ_h² )        # house effect, partially pooled
log τ_p  ~ N( μ_τ[ mode_p , transparency_p ] , σ_τ² )    # excess variance, partially pooled

A pollster with 3 polls is shrunk hard toward its mode/partisanship group mean and retains wide uncertainty; a pollster with 300 polls is largely free. Pollster quality is itself a distribution, and that distribution propagates into every forecast.

6.1 Accuracy vs. lean are different things

h_p (systematic lean) and τ_p (unpredictable noise) are separate parameters on purpose. A pollster with a persistent +2.5 R lean and tiny τ_p is more useful than an unbiased pollster with huge τ_p, because the lean is correctable and the noise is not. Any ranking the system publishes must show both.

6.2 Open empirical questions, to be answered by backtest, not assumption


7. Fundamentals

A separate model producing a prior on α_r before polls arrive:

α_r ~ N( X_r' γ , σ_fund² )

Candidate features (each a hypothesis, not an entitlement): partisan baseline of the geography (§8), incumbency and its interaction with tenure, open-seat status, previous margin, presidential approval, midterm penalty for the president's party, economic indicators as vintage-correct series (§ BACKTESTING FM-03), candidate quality proxies (prior elected office), fundraising (§ build spec 50 — with reverse-causality controls), special-election swing (build spec 49), primary turnout differential, registration trends, and redistricting status.

Feature admission rule. A feature enters only if, in leave-one-cycle-out walk-forward evaluation, it improves out-of-sample log score by more than its standard error, across more than one held-out cycle, with a sign consistent with its hypothesised mechanism. Features that pass in-sample and fail out-of-sample are recorded as negative results in the experiment registry and not silently retried.

With ~35 Senate races per cycle and ~10 usable cycles, the fundamentals design matrix has on the order of 300–400 rows. This supports perhaps 5–10 real parameters, not 30. Regularised regression (horseshoe or hierarchical ridge) is mandatory; stepwise selection is forbidden.

7.1 What is actually fitted, and what was refused

In production: the seat's previous margin in the same seat, widened by measured volatility, with the source cycle's own reference level subtracted and an estimate for the target cycle added back. That de-tiding is a fix under ADR-0017 -- a 2024 prior drawn from 2018 otherwise carries 2018's national environment in full -- and its measured effect is small and honestly reported: 14.26 points of mean absolute error to 14.03, negative in 4 of 11 cycles. For a redistricted House seat, the district's own presidential margin under the new lines replaces the old margin entirely (BaselineRelation, MAE 3.33 against 10.39).

Built, measured, and not in production: elab.models.fundamentals.lean fits, walk-forward,

lean_r = a + b1 * presidential_lean_r + b2 * previous_lean_r + b3 * held_r + b4 * incumbency_r

where a lean is a margin net of its own cycle's reference mean for that office, so the relation describes how a seat sits against the country and says nothing about where the country will sit -- that is the national term, supplied separately and carrying its own error. The recency half-life on training cycles is selected per target cycle by leave-one-cycle-out within the training cycles, because the presidential coefficient for the Senate rises from 0.29 on early cycles to 0.61 by 2024 as ticket splitting collapses, and a fixed half-life would be an assertion about how fast that moves.

At the prior level it is much better: 14.26 points of error against 10.58 over 566 walk-forward seat-cycles, better in 10 of 11 cycles, with standardised errors of sd 0.979 and coverage 0.571 / 0.813 / 0.943 against a nominal 0.50 / 0.80 / 0.95.

It is not in production because it did not clear the gate. On races with 1-3 polls its CRPS gain was 0.053 against a pre-registered threshold of 0.073 (c17_lean_prior, refused), because a poll dominates the blend even after FM-104's correction. On races with no polls -- where the prior is the whole forecast -- it won both folds it could reach by 1.66 and 3.88 points of CRPS, and the pre-registration fixed three; 2004's drift law is unfittable and no unused cycle remains, so it was recorded inconclusive (c18_lean_prior_unpolled). Under ADR-0016 there is no cycle left to register this mechanism against, which makes "cannot be promoted on the evidence available" the answer rather than a pending question. experiments/lean_prior.md and experiments/preregistration_unpolled_prior.md carry both.

7.2 The prior for a race with one to three polls

A race below the fitting threshold is the prior and the polls combined by precision (c7_sparse_blend, promoted). The polls' precision comes from the realised error of a sparse-race poll average, pooled walk-forward from strictly earlier cycles, not from their sampling error: measured over 11 cycles the realised figure is 8.92 points against an assumed 4.04, so weighting by sampling variance gave these polls about five times the weight the evidence supports and quoted an interval narrower by a factor of two. A cycle with nothing behind the measurement falls back to sampling error rather than borrowing a number from its own future, and SparseEstimate.basis records which was used (FM-104, experiments/sparse_blend_poll_weight.md).

7.1 Blending fundamentals with polls

Not an ad-hoc weighted average. The fundamentals distribution is the prior on α_r; polls update it through the likelihood. The effective weight on fundamentals therefore declines automatically as polls accumulate — no hand-tuned decay schedule from fundamentals to polls is needed, and none will be introduced.


8. Geographic baselines

A "partisan baseline" for a state or district must be constructed from results that are comparable across boundary changes.


9. Likely voters, registered voters, undecideds, turnout


10. Correlated error — the term that dominates

Election outcomes are not independent Bernoulli draws. When polling misses, it misses many places at once, in the same direction. The joint error vector e over races at Election Day is:

e  =  σ_nat · 1 · z_0  +  Σ_k σ_k · L_k z_k  +  diag(σ_idio) z_idio
z_*  ~  multivariate Student-t_ν  (not Gaussian)

A factor structure, not a hand-written correlation matrix. Candidate factors, each loading races by an observable characteristic:

Factor Loading variable
National 1 for all races
Region census region / division indicator
Urbanicity share of population in urban/suburban/rural categories
Education share of white population without a bachelor's degree
Race/ethnicity aggregate Black / Hispanic / Asian population shares
Past politics previous presidential margin (continuous loading)
Office type Senate / House / Governor / President
Methodology mix share of the race's polling done online vs. by phone
Turnout environment midterm vs presidential, competitive vs not

σ_nat and the σ_k are estimated from the historical joint distribution of polling errors across races within cycles — i.e. from the empirical fact that 2016, 2020, and 2022 errors were strongly spatially structured. They are not assumed.

Heavy tails are mandatory. ν is estimated, with a prior that permits ν in roughly the 4–20 range. A Gaussian tail assumption would have made the 2016 and 2020 polling misses near-impossible events; history says they are not.

10.1 Sanity constraint

The resulting Σ must reproduce, in posterior predictive checks, the observed cross-race error correlation in held-out cycles. If simulated cycles show less correlated error than history does, the model is overconfident about chamber control even when individual races look fine. This check is a gate, not a nicety (FM-11).


11. The variance budget: nowcast vs forecast

11.0 What the nowcast actually estimates (revised after adversarial review, AR-01)

"If voting occurred today" is not a well-defined quantity until the electorate is specified. Two different questions hide inside the phrase:

Estimand Meaning
(a) The election held today, voted by whoever would actually turn out today — a low-turnout, high-engagement electorate
(b) The election held today, voted by the electorate we project for Election Day

They differ by several points in a midterm. This system estimates (b).

The reason is evidentiary, not aesthetic: likely-voter screens are pollsters' attempts to model the November electorate, so (b) is the quantity the poll data is actually evidence about. Estimating (a) would require turnout models for a hypothetical September election that no data speaks to. (a) is a legitimate question; it is not the one answered here, and the methodology page says so.

Formally: the nowcast marginalises over measurement and composition uncertainty about the projected electorate, but conditions on latent opinion as of today, with no future drift.

For race r at forecast date X, on the latent scale:

NOWCAST:
Var_now = Var[φ_r(X) | D_X]          # state uncertainty: polls, house effects, sparsity
        + σ²_nat                      # industry-wide error shared across races (§10)
        + σ²_race                     # race-specific residual polling error (§5, §10)

FORECAST:
Var_fc  = Var_now  +  Var[ ∫_X^0 dφ_r(t) ]       # integrated future drift

Both halves of the variance-components model, and this is where the model's largest defect lived. §10 decomposes the error of a polling average into mu_c ~ N(0, σ²_nat), one draw per cycle shared by every race in it, and eps_{r,c} ~ N(0, σ²_race), the race-specific residual. On thirteen cycles σ_nat is 2.48 points and σ_race is 5.03 — the residual is the larger term. Until 2026-09-26 the race-level budget carried σ²_nat and, where σ²_race belonged, three unfitted placeholders totalling 1.89 points² against the 25.3 the fitted term carries (σ²_lv 1.00, σ²_und 0.64, σ²_model 0.25). The chamber simulator had been using σ_race all along.

Restoring it, walk-forward, took realised coverage on 247 scored races from 0.259 / 0.518 / 0.680 to 0.474 / 0.777 / 0.895 against nominal 0.50 / 0.80 / 0.90, with nothing tuned (FM-79). The three placeholders are replaced rather than joined, because a residual already contains the likely-voter, undecided and specification error they stood for; carrying both over-widens. The budget therefore contains no unfitted term, which is why allow_guesses no longer appears on the publishing path — the refusal in §11's own implementation now guards the published forecast instead of being waived by every caller.

The independent check that σ_race is the right term rather than a well-sized fudge: realised error with each cycle's mean removed — which is what eps is — has an rms of 5.42 against a fitted 5.03.

The drift term is estimated empirically: for each historical cycle, measure the realised change in the latent state (or, failing that, in a consistently-computed polling average) between horizon h and Election Day, pooled across races, and fit Var_drift(h). A pure random walk implies Var ∝ h; history may show mean reversion (sublinear) or campaign-event clustering (superlinear, with conventions and debates). We fit it; we do not assume ∝ h.

Invariants enforced by test:

This is the structural reason a race can be "D+6 today with an 88% chance if the election were today" and simultaneously "68% to win in November". Both numbers are always published.

11.1 The chamber products need their own drift, and the House's is national

The budget above is a race's. A chamber forecast built from priors rather than from race-level polling has a different structure and, until 2026-09-25, was missing this term entirely.

The House is 435 districts of which a handful are ever polled, so its forecast is each district's prior shifted by the change in the national environment and spread by that district's volatility net of the nation. The national shift is measured from current generic-ballot polls. What was absent was any term for the movement of that environment between the forecast date and Election Day: the district spread is an election-to-election quantity, and σ_nat is the cycle-level polling error, so a forecast made 41 days out carried the same national uncertainty as one made on the morning of the election. That is why P(Democratic majority) printed as 0.995.

elab.models.fundamentals.environment_drift fits Var_env(h) = a·h^b from non-overlapping increments of the generic-ballot path, four cycles of it, and the forecast adds Var_env(h) to σ²_nat in quadrature. Three properties of the estimate are worth stating because they are easy to get wrong:

And a third national term: how well today is measured. Var_env(h) is the movement still to come; it says nothing about the precision of the starting point. Until 2026-09-25 the budget treated a window of 9 polls from 5 firms as exactly as precise as one of 165 polls from 43 — the estimate's own error appeared nowhere. It is now added in quadrature as max of two figures, neither of which bounds the other (FM-74, FM-75, FM-76):

1e4/independent_n — the arithmetic this term started as — is now a printed diagnostic only. It is blind to both dominant components: it said 0.63 points for the 2020 window where the dispersion law says 1.84, and 0.63 for 2022 where the jackknife says 2.36, because Morning Consult held 54% of that window's weight at D+3.19 while one 108,206-respondent SurveyMonkey panel held 20% at D−2.00. Each route wins somewhere — the law in 2020, 2024 and the live 2026 window, the jackknife in 2022 — which is why the term takes the larger rather than choosing between them.

The term is what makes an early, thinly polled or concentrated environment print a wider chamber interval than a late, broad one: for 2026 it took the House seat sd from 16.5 to 18.0 and P(D majority) from 0.994 to 0.991.

It double-counts whatever share of σ_nat is the generic ballot's own sampling noise. That share is bounded by the same measurement: the cycles σ_nat is estimated from carried 0.6–0.9 points of it, which is 0.4 of about 10 points², a 2% overstatement of the national sd in a well-polled window and inside the resolution of an 11-cycle estimate. The direction is honest widening, and the alternative in a thin window is claiming a precision the polls do not have.

11.1.1 Two effective sample sizes for the environment, always reported together

The national environment is a recency- and precision-weighted mean of generic-ballot polls, and how much information it rests on has two answers that differ by more than an order of magnitude:

In the 2020 window these are 695,887 and 26,032: a sampling sd for the margin of 0.12 points against 0.62. describe() prints both and sampling_sd quotes the clustered one. Neither is available alone, because the gap between them is the substance. estimate() needs series_window( labelled=True) to compute the second at all, and returns None for it rather than a guess when the pollster is not supplied.

The estimate is insensitive to this — one vote per firm instead of one vote per poll changes the House seat error not at all (experiments/environment_weighting.md) — so what was wrong was the advertised precision, not the number.

11.1.2 What a poll's precision is actually made of

Both sample sizes above assume a poll's variance is proportional to 1/n. Measured on 13,639 pairs of generic-ballot polls fielded within three days of each other by different firms, with each firm's house effect subtracted:

Var(A − B) = 4.04 + 8,005 · (1/n_A + 1/n_B)

Pairs rather than deviations from a consensus, because a consensus carries an error of its own that is common to every poll at that date and would land in the intercept whether or not the polls were clean.

The house-mix term dominates the environment's variance in every window measured. The implication for the weights — that a 108,206-respondent panel is worth 4.1 times a 1,000-person poll rather than 108 times — was registered as c11_inverse_variance_weights and is applied as of 2026-09-27 (house_prior/0.9.0). Promoted on its pre-registered rule: a firm-clustered 95% interval of [+0.063, +0.329] pts² on 3,765 held-out polls over 127 firms, excluding zero, reproducing the [+0.059, +0.304] measured at registration. The metric is leave-one-poll-out against other firms' polls in a 21-day window, so it reads no election outcome and consumed no holdout under ADR-0010. experiments/preregistration_inverse_variance.md has the decision; the environment moved by about a third of a House seat, which is what the registration predicted and is not the argument for it.

11.2 What is shared across races, and where it comes from

For a chamber or the electoral college, the question "which part of a race's uncertainty moves with every other race" decides the answer, and the budget above already names it: σ²_syst is industry-wide polling error, one number per cycle, identical across races by construction. Everything else in a race's budget is that race's own.

So the electoral-college simulation draws σ_syst once per simulated election from the stored budget of the races it is combining, and draws every other term per state. It does not draw the election-to-election national swing on top of states whose margins have already been measured from polls — which both earlier implementations did, with the effect that a 2024 forecast came out P(D wins) = 0.585 with a 90% interval 140 electors wide in a year decided by a point and a half (FM-66). Unpolled states take the prior shifted by the swing the polled states imply, plus the idiosyncratic part of a state's swing and the standard error of that swing estimate.


12. Simulation

  1. Draw S posterior samples of all parameters (S set by MCSE targets, §15.3, not by round numbers).
  2. For each draw, draw the correlated error vector e (§10) and, for forecasts, the drift.
  3. Transform from ALR space back to shares; compute margins; determine per-race winners.
  4. Aggregate: House seats, Senate seats (including holdover seats, independents' caucusing, and VP tie-break), Governors, Electoral College, national popular vote.
  5. Persist the full joint draw matrix, not just summaries — chamber-control probability is a functional of the joint distribution and cannot be recovered from marginals.

Convergence is checked by Monte Carlo standard error on the reported quantities, and the simulation seed is recorded (N-01).


13. Baselines the model must beat

Every sophisticated component is evaluated against, at minimum:

Baseline Description
B0 Previous election result in the same geography
B1 Geography partisan baseline (§8), uniform national swing applied
B2 Unweighted mean of last 21 days of polls
B3 Exponentially recency-weighted poll average, decay fitted by grid search
B4 B3 + sample-size weighting + house-effect correction
B5 Fundamentals-only (§7)
B6 Simple 50/50 blend of B4 and B5

Plus an uncertainty baseline: B2 with an empirical error distribution taken straight from the historical distribution of (poll average − result). This is a surprisingly strong probabilistic baseline and beating it is the real test.

If the hierarchical Bayesian model cannot beat B4/B6 out of sample on log score and calibration, the simple model ships and the complex one goes back to the lab. This is a rule, not a sentiment.


14. Model uncertainty

Parameter uncertainty is not the same as being right about the model. We carry σ²_model by:

  1. Fitting a small set of structurally distinct specifications (different decay families, different correlation factor sets, different observation distributions).
  2. Evaluating each out of sample.
  3. Combining by stacking weights fitted on held-out log score — not by equal weights and not by in-sample marginal likelihood (which rewards overfitting under misspecification).
  4. Reporting the between-specification spread as an explicit variance component.

Stacking must itself be validated: an ensemble is a hypothesis too (build spec 27).


15. Estimation and diagnostics

15.1 Engine

CPU NUTS. Primary candidate: PyMC with the nutpie sampler. Fallback: CmdStan via cmdstanpy. Selection by benchmark on the Phase 1 race (ADR-0004): ESS per wall-clock second, divergence behaviour, and memory footprint under 16 threads. No framework is adopted for fashion.

15.2 Priors

Every prior is documented with its justification in a table in this file as it is implemented, and every prior is subjected to a prior predictive check: simulate from the prior alone and confirm the implied elections are politically plausible (no 97%-of-vote Senate winners, no 40-point national swings in a week).

15.3 Diagnostic gates (a fit that fails these is not used)

The gate inspects every sampled parameter, not a chosen few. An early implementation summarised four hyperparameters and reported max_rhat = 1.0017 while the sampler itself was warning that R-hat exceeded 1.01 elsewhere. That is worse than a weaker gate: it produces a single number that reads like a guarantee over the whole model. The diagnostics now record which parameter is worst, so a failure names its cause.


16. Known weaknesses, stated up front

  1. Few effective observations of systematic polling error (§5). Roughly 10–20 cycle-office units. σ_nat is weakly identified and any claim to know it precisely is false.
  2. Non-stationarity. Polling methodology changed radically between 2000 and 2026 (landline → online panel). Historical error estimates may not transfer. Mitigation: concept-drift weighting, tested out of sample (build spec 82) — not a solution, a partial mitigation.
  3. House data sparsity. Most of the 435 districts are never polled. Their forecasts are fundamentals-plus-national-swing, and their uncertainty must reflect that. The chamber distribution is therefore driven largely by the correlated error structure, which is the least well-identified part of the model.
  4. Candidate-level idiosyncrasy (scandals, gaffes, health events) is unmodellable prospectively and enters only as idiosyncratic variance.
  5. Selection in the poll set. Polls are released, not sampled. Suppressed unfavourable internal polls bias the observed set in ways no weighting fixes (FM-08).
  6. Backtest reuse. Every pass over the same 10 cycles burns holdout. The holdout ledger (BACKTESTING.md §6) tracks the burn rate; when it is exhausted, honest evaluation requires waiting for a real election.