The statistical model
Every distribution, every prior and every term in the variance budget, with the assumption each one rests on written next to it.
STATISTICAL MODEL
This document specifies the mathematics. It is written for a statistician who has not read the code and who intends to find the mistakes. Every constant that appears below is either estimated from data, read from configuration, or accompanied by a justification. There are no unexplained magic numbers.
Status, 2026-09-27: fitted, and validated where it says so. This line read "design
specification, not yet fitted. Nothing here is validated" until today, which was written before
anything was fitted and left standing for months afterwards — so the project's primary statistical
document asserted the opposite of the truth. The model has been fitted, walk-forward backtested and
scored on 435 races over ten cycles, 2006-2024 (single_race_latent/0.7.0: MAE 4.89, lean +2.08,
realised coverage 0.471 / 0.798 / 0.903 against a nominal 0.50 / 0.80 / 0.90 -- all three inside the
ADR-0018 floor), and two challengers have been promoted through a pre-registered gate. On those same
races a recency- and size-weighted polling average scores MAE 5.51 with a lean of +2.86, so the model
beats its baseline and removes 0.77 points of the polls' lean rather than adding to it
(F-11, experiments/calibration_decomposition.md).
The five-cycle figures this paragraph used to quote (295 races, MAE 4.68, coverage
0.475 / 0.810 / 0.919) were for 0.6.0 before the evidence base was extended to 2006 and before
FM-104's fix; they are not comparable to the ten-cycle ones and are not the same population.
What remains genuinely unvalidated is named where it arises, not hedged here. The cross-race posterior
predictive check now passes on all six cycles, 2014 at tail p = 0.089 where it failed at 0.037
(experiments/phase5_validation_2026-09-29.json, experiments/phase5_exit_met.md) -- a pass, and the
tightest of the six against a next-closest 0.110. sigma_region fits at 0.00, and that is now a bounded claim rather than a shrug: measured on the real skeleton, the estimator finds a true 2.0-point regional term 98% of the time and a 1.0-point one 38% of the time, so a regional term of 2.0 points or more is ruled out and one of 1.0 or less is undetectable here. The forecast's own regional signal, debiased for the noise on each regional mean, is 0.72 points across all races and 0.00 among well-polled ones — below the floor, so the fitted zero and the slice's 3.33-point spread agree. A spread over four groups is about twice a standard deviation and each regional mean carries ±1.2 of noise, which is where the apparent contradiction came from (experiments/region_tension.md). sigma_office fits at 0.00 on the same estimator and its floor also lands at 2.0 points, but more weakly — 82% detection at a true 2.0 against the regional 98% — so it bounds less. Where a parameterisation below is still a competing
candidate rather than a decision, it says so in place.
1. Notation
| Symbol | Meaning |
|---|---|
r |
race (a contest in one geography for one office in one cycle) |
c |
candidate within a race; K_r candidates plus an "undecided/other" category |
t |
time, in days, with t = 0 at Election Day and t < 0 before it |
i |
an individual poll observation |
p(i) |
pollster of poll i; f(p) its pollster family |
n_i |
nominal sample size |
y_i |
reported vector of candidate shares |
θ_r(t) |
latent population support in race r at time t |
D_X |
the information set knowable at calendar date X |
Two quantities of interest, always distinct:
- Nowcast at date
X: the distribution of the result of an election held on X, givenD_X. - Forecast at date
X: the distribution of the result of the election held on Election Day, givenD_X.
They differ by exactly one term: integrated future opinion drift. §11 makes that precise.
2. Scale
Vote shares are compositional; modelling them on the raw simplex with Gaussian noise is wrong at the boundaries and produces impossible intervals in lopsided races.
We work in additive log-ratio (ALR) space:
φ_rc(t) = log( θ_rc(t) / θ_r,ref(t) ), c = 1 … K_r - 1
Choice of reference category (revised after adversarial review, AR-10). The reference sits in the denominator of every log-ratio, so it must not approach zero. Undecided is therefore a bad choice: it shrinks toward zero as Election Day nears, and the log-ratios would destabilise precisely in the final fortnight when precision matters most.
The reference is the leading major-party candidate, fixed per race at fit time from the fundamentals prior — chosen from the prior, not from the data being fitted, so the parameterisation does not depend on the likelihood. Undecided remains an explicit modelled category (§9); it is simply not the denominator. The centred-log-ratio (CLR) parameterisation, which avoids choosing a reference at all, is retained as the tested alternative.
For the common two-candidate case this reduces to a single logit of the two-party share, and
the margin m_r = θ_rA - θ_rB is recovered by the inverse transform inside the simulation, so
all uncertainty propagates to the margin without a normal approximation on the margin itself.
Candidate alternative to test: a Dirichlet-multinomial observation model directly on the simplex. Selection is empirical (ADR-0005), not aesthetic.
3. Latent opinion: the state process
For each race, the latent vector evolves as a Gaussian random walk with a national component:
φ_r(t) = α_r + β_r · η(t) + u_r(t)
-
α_r— race-level intercept. Its prior comes from the fundamentals model (§7). -
η(t)— national latent environment, a scalar (or low-dimensional) random walk shared across all races in the cycle:η(t) = η(t-1) + ν_t,ν_t ~ N(0, σ_η²). -
β_r— race elasticity: how much racermoves when the nation moves 1 unit. It is not fixed at 1: Senate races with strong incumbents historically move less than the nation; open presidential-year races move more.Identification (revised after adversarial review, AR-03).
α_randβ_rare not jointly identified from a single cycle's polls — any level shift can be absorbed by either. Thereforeβ_ris estimated out of cycle, from hierarchically-shrunk regressions of that geography's swing on national swing across prior cycles, and enters the live fit as a fixed informative prior carrying its historical uncertainty. It is never a free parameter in the within-cycle fit. Like everything else, it is re-estimated at each backtest as-of date from data available then. -
u_r(t)— race-idiosyncratic random walk,u_r(t) = u_r(t-1) + ω_t,ω_t ~ N(0, σ_u,r²).
σ_η and σ_u are not chosen by hand. They are estimated from historical polling
trajectories: the model is fitted on past cycles where the truth is known, and the innovation
scales are hyperparameters with weakly-informative priors updated by that history. The
posterior for σ_η is then used as a prior for the live cycle.
3.1 Why a state-space model rather than a weighted average
A weighted average answers "what is the centre of recent polls". A state-space model answers
"what is public opinion, given that polls are noisy, biased, and irregularly spaced". The
second is the question we actually have. Critically, it borrows strength across time and
across races through η(t), so a state with two polls is informed by the national trend rather
than by its own two polls alone.
This must still be proved to beat a recency-weighted average out of sample (§13). If it does not, we ship the average.
4. Observation model for polls
For poll i fielded over [t_start, t_end] in race r:
ỹ_i ~ Student-t_ν ( μ_i , Σ_i )
μ_i = φ_r(t̄_i) + h_{p(i)} + m_{mode(i)} + g_{pop(i)}(t̄_i) + s_i · ψ
where
t̄_iis the effective field midpoint, weighted by field-period length. A poll fielded over 10 days is not a measurement of its last day.h_p— pollster house effect (§5, §6).m_mode— polling-mode effect (live caller / IVR / online panel / text / mixed).g_pop(t)— population effect for adults vs registered voters vs likely voters, allowed to vary withtso that the LV/RV gap can shrink or grow across the calendar (§9).s_i ∈ {-1,0,+1}— partisan sponsorship indicator;ψthe estimated sponsorship shift.
4.1 Variance
Σ_i = DEFF_i · Σ_sampling(n_i, θ) + τ²_{p(i)} + τ²_{family(p(i))}
Σ_samplingis the multinomial sampling covariance implied byn_i. It is a floor, not the truth: reported margins of error systematically understate real error because weighting, clustering, and nonresponse inflate variance.DEFF_i— design effect. Estimated per mode and per pollster where data allow; where not, a hierarchical prior centred on a value estimated from the historical error distribution. We do not assumeDEFF = 1.τ²_p— pollster excess variance beyond sampling (§6).- Student-t with estimated
νrather than Gaussian: individual polls produce outliers far more often than a normal allows.
4.2 Non-independence
Three mechanisms break the "one poll = one independent observation" assumption, and each is modelled explicitly rather than ignored:
- Tracking polls / rolling samples. Consecutive releases share respondents. Overlapping
releases from the same tracker are collapsed into non-overlapping windows, or given a
correlation term
ρ_overlapproportional to the fraction of shared field days. - Shared samples across questions. One survey reporting Senate and Governor toplines is one sample; its errors correlate across races. Modelled with a per-sample random effect shared by all races it reports.
- Pollster families. Firms sharing panels, vendors, or methodology (
f(p)) get a family-level house effect and family-level variance term, so releasing five polls under five brand names does not buy five independent observations.
5. Identification — the part most implementations get wrong
The decomposition μ_i = φ_r(t) + h_p is not identified without a constraint: adding a
constant to every h_p and subtracting it from every φ_r leaves the likelihood unchanged.
Polls alone can never tell us the absolute level of opinion. Only election results can. Therefore:
- Within a cycle, house effects are identified only relative to each other. We impose a
volume-weighted sum-to-zero constraint on
h_pwithin each cycle so the latent state is pinned to "the average pollster". - The industry-wide bias — the amount by which the average pollster misses the result — is
a separate parameter
b_cycle, estimable only from realised results across past cycles. It is the dominant term in §10's correlated error. - Consequently, a model fitted on polls alone and reporting tight intervals is reporting the uncertainty of "where the polling average is", not "where the election lands". Conflating these is the single most common failure in public election models. The variance budget in §11 keeps them apart.
Number of effective observations of industry-wide bias is small — roughly one per
office-type per cycle, so on the order of 10–20 for the modern polling era. b_cycle's scale
parameter is therefore weakly identified and must carry a wide, honestly-reported posterior.
This is a fundamental limit of the problem, not of the implementation, and is recorded as
failure mode FM-14.
6. Pollster quality, hierarchically
No subjective letter grades. For each pollster p:
h_p ~ N( μ_h[ mode_p , partisan_p ] , σ_h² ) # house effect, partially pooled
log τ_p ~ N( μ_τ[ mode_p , transparency_p ] , σ_τ² ) # excess variance, partially pooled
A pollster with 3 polls is shrunk hard toward its mode/partisanship group mean and retains wide uncertainty; a pollster with 300 polls is largely free. Pollster quality is itself a distribution, and that distribution propagates into every forecast.
6.1 Accuracy vs. lean are different things
h_p (systematic lean) and τ_p (unpredictable noise) are separate parameters on purpose. A
pollster with a persistent +2.5 R lean and tiny τ_p is more useful than an unbiased pollster
with huge τ_p, because the lean is correctable and the noise is not. Any ranking the system
publishes must show both.
6.2 Open empirical questions, to be answered by backtest, not assumption
- Does
h_ppersist across cycles, and with what autocorrelation? Model ash_{p,cycle} = λ · h_{p,cycle-1} + e, estimateλ. Ifλ ≈ 0, historical house effects are useless for prediction and must not be applied. - Do pollsters improve or deteriorate? Test a trend term.
- Does a methodology change reset a pollster's history? Treat a documented methodology change
as a partial break: variance inflation on the carried-forward
h_p. - Survivorship: pollsters that quit after a bad cycle are missing from the evaluation set. Their absence biases estimated industry accuracy optimistically (FM-07).
7. Fundamentals
A separate model producing a prior on α_r before polls arrive:
α_r ~ N( X_r' γ , σ_fund² )
Candidate features (each a hypothesis, not an entitlement): partisan baseline of the geography (§8), incumbency and its interaction with tenure, open-seat status, previous margin, presidential approval, midterm penalty for the president's party, economic indicators as vintage-correct series (§ BACKTESTING FM-03), candidate quality proxies (prior elected office), fundraising (§ build spec 50 — with reverse-causality controls), special-election swing (build spec 49), primary turnout differential, registration trends, and redistricting status.
Feature admission rule. A feature enters only if, in leave-one-cycle-out walk-forward evaluation, it improves out-of-sample log score by more than its standard error, across more than one held-out cycle, with a sign consistent with its hypothesised mechanism. Features that pass in-sample and fail out-of-sample are recorded as negative results in the experiment registry and not silently retried.
With ~35 Senate races per cycle and ~10 usable cycles, the fundamentals design matrix has on the order of 300–400 rows. This supports perhaps 5–10 real parameters, not 30. Regularised regression (horseshoe or hierarchical ridge) is mandatory; stepwise selection is forbidden.
7.1 What is actually fitted, and what was refused
In production: the seat's previous margin in the same seat, widened by measured volatility, with
the source cycle's own reference level subtracted and an estimate for the target cycle added back.
That de-tiding is a fix under ADR-0017 -- a 2024 prior drawn from 2018 otherwise carries 2018's
national environment in full -- and its measured effect is small and honestly reported: 14.26 points
of mean absolute error to 14.03, negative in 4 of 11 cycles. For a redistricted House seat, the
district's own presidential margin under the new lines replaces the old margin entirely
(BaselineRelation, MAE 3.33 against 10.39).
Built, measured, and not in production: elab.models.fundamentals.lean fits, walk-forward,
lean_r = a + b1 * presidential_lean_r + b2 * previous_lean_r + b3 * held_r + b4 * incumbency_r
where a lean is a margin net of its own cycle's reference mean for that office, so the relation describes how a seat sits against the country and says nothing about where the country will sit -- that is the national term, supplied separately and carrying its own error. The recency half-life on training cycles is selected per target cycle by leave-one-cycle-out within the training cycles, because the presidential coefficient for the Senate rises from 0.29 on early cycles to 0.61 by 2024 as ticket splitting collapses, and a fixed half-life would be an assertion about how fast that moves.
At the prior level it is much better: 14.26 points of error against 10.58 over 566 walk-forward seat-cycles, better in 10 of 11 cycles, with standardised errors of sd 0.979 and coverage 0.571 / 0.813 / 0.943 against a nominal 0.50 / 0.80 / 0.95.
It is not in production because it did not clear the gate. On races with 1-3 polls its CRPS gain was
0.053 against a pre-registered threshold of 0.073 (c17_lean_prior, refused), because a poll
dominates the blend even after FM-104's correction. On races with no polls -- where the prior is
the whole forecast -- it won both folds it could reach by 1.66 and 3.88 points of CRPS, and the
pre-registration fixed three; 2004's drift law is unfittable and no unused cycle remains, so it was
recorded inconclusive (c18_lean_prior_unpolled). Under ADR-0016 there is no cycle left to register
this mechanism against, which makes "cannot be promoted on the evidence available" the answer rather
than a pending question. experiments/lean_prior.md and
experiments/preregistration_unpolled_prior.md carry both.
7.2 The prior for a race with one to three polls
A race below the fitting threshold is the prior and the polls combined by precision (c7_sparse_blend,
promoted). The polls' precision comes from the realised error of a sparse-race poll average,
pooled walk-forward from strictly earlier cycles, not from their sampling error: measured over 11
cycles the realised figure is 8.92 points against an assumed 4.04, so weighting by sampling variance
gave these polls about five times the weight the evidence supports and quoted an interval narrower by a
factor of two. A cycle with nothing behind the measurement falls back to sampling error rather than
borrowing a number from its own future, and SparseEstimate.basis records which was used
(FM-104, experiments/sparse_blend_poll_weight.md).
7.1 Blending fundamentals with polls
Not an ad-hoc weighted average. The fundamentals distribution is the prior on α_r; polls
update it through the likelihood. The effective weight on fundamentals therefore declines
automatically as polls accumulate — no hand-tuned decay schedule from fundamentals to polls is
needed, and none will be introduced.
8. Geographic baselines
A "partisan baseline" for a state or district must be constructed from results that are comparable across boundary changes.
- Districts are identified by
(cycle, state, district_number)plus a boundary hash. A district whose lines changed is a different unit even at the same number. - Baselines for redrawn districts are rebuilt from precinct-level results disaggregated onto new boundaries where precinct data exist; where they do not, the baseline carries an explicit larger uncertainty rather than a silent guess.
- Uncontested races produce missing, not zero, two-party margins; they are imputed with uncertainty from neighbouring-office results, and flagged.
- Baselines are a weighted blend of recent presidential results and down-ballot results, with the weights estimated, and with a drift term for realignment (build spec 82).
9. Likely voters, registered voters, undecideds, turnout
g_pop(t): the LV-vs-RV effect is estimated, allowed to vary by cycle, office, and calendar time. The folk belief of a fixed "LV premium" of a couple of points for one party is a hypothesis to be tested per era, not a constant.- Undecideds are a modelled category, not a rounding error. Their allocation at Election Day is a parameter with a prior estimated from historical late-break behaviour, allowed to depend on incumbency (the "incumbent rule" is testable and, in recent data, weak), the presence of third-party candidates, and the size of the undecided bloc itself.
- Turnout composition error enters as a distinct variance component, because pollsters' likely-voter screens are themselves models that can fail in a correlated way across the industry. It is not folded into sampling error.
10. Correlated error — the term that dominates
Election outcomes are not independent Bernoulli draws. When polling misses, it misses many
places at once, in the same direction. The joint error vector e over races at Election Day is:
e = σ_nat · 1 · z_0 + Σ_k σ_k · L_k z_k + diag(σ_idio) z_idio
z_* ~ multivariate Student-t_ν (not Gaussian)
A factor structure, not a hand-written correlation matrix. Candidate factors, each loading races by an observable characteristic:
| Factor | Loading variable |
|---|---|
| National | 1 for all races |
| Region | census region / division indicator |
| Urbanicity | share of population in urban/suburban/rural categories |
| Education | share of white population without a bachelor's degree |
| Race/ethnicity | aggregate Black / Hispanic / Asian population shares |
| Past politics | previous presidential margin (continuous loading) |
| Office type | Senate / House / Governor / President |
| Methodology mix | share of the race's polling done online vs. by phone |
| Turnout environment | midterm vs presidential, competitive vs not |
σ_nat and the σ_k are estimated from the historical joint distribution of polling errors
across races within cycles — i.e. from the empirical fact that 2016, 2020, and 2022 errors were
strongly spatially structured. They are not assumed.
Heavy tails are mandatory. ν is estimated, with a prior that permits ν in roughly the
4–20 range. A Gaussian tail assumption would have made the 2016 and 2020 polling misses
near-impossible events; history says they are not.
10.1 Sanity constraint
The resulting Σ must reproduce, in posterior predictive checks, the observed cross-race error correlation in held-out cycles. If simulated cycles show less correlated error than history does, the model is overconfident about chamber control even when individual races look fine. This check is a gate, not a nicety (FM-11).
11. The variance budget: nowcast vs forecast
11.0 What the nowcast actually estimates (revised after adversarial review, AR-01)
"If voting occurred today" is not a well-defined quantity until the electorate is specified. Two different questions hide inside the phrase:
| Estimand | Meaning |
|---|---|
| (a) | The election held today, voted by whoever would actually turn out today — a low-turnout, high-engagement electorate |
| (b) | The election held today, voted by the electorate we project for Election Day |
They differ by several points in a midterm. This system estimates (b).
The reason is evidentiary, not aesthetic: likely-voter screens are pollsters' attempts to model the November electorate, so (b) is the quantity the poll data is actually evidence about. Estimating (a) would require turnout models for a hypothetical September election that no data speaks to. (a) is a legitimate question; it is not the one answered here, and the methodology page says so.
Formally: the nowcast marginalises over measurement and composition uncertainty about the projected electorate, but conditions on latent opinion as of today, with no future drift.
For race r at forecast date X, on the latent scale:
NOWCAST:
Var_now = Var[φ_r(X) | D_X] # state uncertainty: polls, house effects, sparsity
+ σ²_nat # industry-wide error shared across races (§10)
+ σ²_race # race-specific residual polling error (§5, §10)
FORECAST:
Var_fc = Var_now + Var[ ∫_X^0 dφ_r(t) ] # integrated future drift
Both halves of the variance-components model, and this is where the model's largest defect lived.
§10 decomposes the error of a polling average into mu_c ~ N(0, σ²_nat), one draw per cycle shared by
every race in it, and eps_{r,c} ~ N(0, σ²_race), the race-specific residual. On thirteen cycles
σ_nat is 2.48 points and σ_race is 5.03 — the residual is the larger term. Until 2026-09-26 the
race-level budget carried σ²_nat and, where σ²_race belonged, three unfitted placeholders totalling
1.89 points² against the 25.3 the fitted term carries (σ²_lv 1.00, σ²_und 0.64, σ²_model 0.25).
The chamber simulator had been using σ_race all along.
Restoring it, walk-forward, took realised coverage on 247 scored races from 0.259 / 0.518 / 0.680 to
0.474 / 0.777 / 0.895 against nominal 0.50 / 0.80 / 0.90, with nothing tuned (FM-79). The three
placeholders are replaced rather than joined, because a residual already contains the likely-voter,
undecided and specification error they stood for; carrying both over-widens. The budget therefore
contains no unfitted term, which is why allow_guesses no longer appears on the publishing path — the
refusal in §11's own implementation now guards the published forecast instead of being waived by every
caller.
The independent check that σ_race is the right term rather than a well-sized fudge: realised error with
each cycle's mean removed — which is what eps is — has an rms of 5.42 against a fitted 5.03.
The drift term is estimated empirically: for each historical cycle, measure the realised
change in the latent state (or, failing that, in a consistently-computed polling average)
between horizon h and Election Day, pooled across races, and fit Var_drift(h). A pure
random walk implies Var ∝ h; history may show mean reversion (sublinear) or campaign-event
clustering (superlinear, with conventions and debates). We fit it; we do not assume ∝ h.
Invariants enforced by test:
Var_fc ≥ Var_nowfor allh > 0.Var_fc → Var_nowash → 0.Var_fcis monotone non-decreasing inhunless a fitted mean-reversion term is justified by held-out evidence and documented.
This is the structural reason a race can be "D+6 today with an 88% chance if the election were today" and simultaneously "68% to win in November". Both numbers are always published.
11.1 The chamber products need their own drift, and the House's is national
The budget above is a race's. A chamber forecast built from priors rather than from race-level polling has a different structure and, until 2026-09-25, was missing this term entirely.
The House is 435 districts of which a handful are ever polled, so its forecast is each district's
prior shifted by the change in the national environment and spread by that district's
volatility net of the nation. The national shift is measured from current generic-ballot polls.
What was absent was any term for the movement of that environment between the forecast date and
Election Day: the district spread is an election-to-election quantity, and σ_nat is the
cycle-level polling error, so a forecast made 41 days out carried the same national uncertainty as
one made on the morning of the election. That is why P(Democratic majority) printed as 0.995.
elab.models.fundamentals.environment_drift fits Var_env(h) = a·h^b from non-overlapping
increments of the generic-ballot path, four cycles of it, and the forecast adds Var_env(h) to
σ²_nat in quadrature. Three properties of the estimate are worth stating because they are easy to
get wrong:
-
The race-level drift law is the wrong instrument here. It is fitted on single races and therefore contains idiosyncratic movement, which is exactly what averages out of a national number. Applying it would overstate the term.
-
Measurement noise is subtracted. Each end of an increment is an estimate of the environment, so an observed difference is movement plus twice that estimate's sampling error. At short horizons the noise is the larger part -- 31% of the raw variance overall. It is also why every horizon exceeds the estimation window: overlapping windows share most of their polls, so their errors cancel rather than add, and the independent-sample correction then subtracts a variance that is not there.
Which effective sample size is a real choice, because the estimate has two (§11.1.1) and the original code used a third quantity that was neither (FM-73). The correction uses
sampling_n, the independence-assuming one. The firm-clustered figure is right for the level and wrong for a difference, and the data settles it rather than the argument: increments whose ends draw 97% of their weight from the same firms have a raw variance of 0.70 pts² at 21 days against a clustered noise estimate of 2.06 pts². A component three times the total it belongs to is refuted. House effects and panel persistence appear at both ends of such a difference and cancel. -
The fits are mean-reverting and the exponent is constrained. On one cycle's path the unconstrained fit gave
b = -0.78, which says a forecast three months out is more certain than one a month out.bis clamped into[0, 1.5]and the law records that it was clamped, so a clamped fit cannot be read as a measured one.
And a third national term: how well today is measured. Var_env(h) is the movement still to
come; it says nothing about the precision of the starting point. Until 2026-09-25 the budget treated
a window of 9 polls from 5 firms as exactly as precise as one of 165 polls from 43 — the estimate's
own error appeared nowhere. It is now added in quadrature as max of two figures, neither of which
bounds the other (FM-74, FM-75, FM-76):
- the measured dispersion law (§11.1.2).
Var(poll) = floor + slope/nwith a house-effect dispersion across firms, fitted on same-week cross-firm pairs and walk-forward per cycle. The weighted mean's variance follows, and it has two parts: a per-poll part, and a house-mix part that depends on how concentrated the firm weights are because a firm's house effect is one number however many polls carry it. - the firm-level jackknife: delete one firm, re-estimate, take the cluster-robust spread. This is what this window's firms disagree about rather than what the corpus average says, and it is withheld below four firms rather than reported from two deviations.
1e4/independent_n — the arithmetic this term started as — is now a printed diagnostic only. It is
blind to both dominant components: it said 0.63 points for the 2020 window where the dispersion law
says 1.84, and 0.63 for 2022 where the jackknife says 2.36, because Morning Consult held 54% of that
window's weight at D+3.19 while one 108,206-respondent SurveyMonkey panel held 20% at D−2.00. Each
route wins somewhere — the law in 2020, 2024 and the live 2026 window, the jackknife in 2022 — which is
why the term takes the larger rather than choosing between them.
The term is what makes an early, thinly polled or concentrated environment print a wider chamber interval than a late, broad one: for 2026 it took the House seat sd from 16.5 to 18.0 and P(D majority) from 0.994 to 0.991.
It double-counts whatever share of σ_nat is the generic ballot's own sampling noise. That share is
bounded by the same measurement: the cycles σ_nat is estimated from carried 0.6–0.9 points of it,
which is 0.4 of about 10 points², a 2% overstatement of the national sd in a well-polled window and
inside the resolution of an 11-cycle estimate. The direction is honest widening, and the alternative
in a thin window is claiming a precision the polls do not have.
11.1.1 Two effective sample sizes for the environment, always reported together
The national environment is a recency- and precision-weighted mean of generic-ballot polls, and how much information it rests on has two answers that differ by more than an order of magnitude:
sampling_n-- Kish's(sum w)² / sum(w²/nᵢ), the size that reproduces the weighted mean's sampling variance if the polls are independent samples. With recency weights this is larger thansum(rᵢ·nᵢ), because down-weighting a three-week-old poll gives up freshness, not respondents.independent_n-- each firm treated as one cluster whose precision is capped at its largest single wave. 59% of this corpus carries 538'strackingflag and 72% of the 2020 window's weight is two firms re-interviewing the same panels weekly, so respondents are being counted repeatedly.
In the 2020 window these are 695,887 and 26,032: a sampling sd for the margin of 0.12 points against
0.62. describe() prints both and sampling_sd quotes the clustered one. Neither is available
alone, because the gap between them is the substance. estimate() needs series_window( labelled=True) to compute the second at all, and returns None for it rather than a guess when the
pollster is not supplied.
The estimate is insensitive to this — one vote per firm instead of one vote per poll changes the
House seat error not at all (experiments/environment_weighting.md) — so what was wrong was the
advertised precision, not the number.
11.1.2 What a poll's precision is actually made of
Both sample sizes above assume a poll's variance is proportional to 1/n. Measured on 13,639 pairs of
generic-ballot polls fielded within three days of each other by different firms, with each firm's
house effect subtracted:
Var(A − B) = 4.04 + 8,005 · (1/n_A + 1/n_B)
Pairs rather than deviations from a consensus, because a consensus carries an error of its own that is common to every poll at that date and would land in the intercept whether or not the polls were clean.
- Sampling behaves about as stated. Slope 8,005 against the 10,000 a simple random sample of a two-party margin implies: a design effect near one.
- Each poll carries 2.02 pts² = 1.42 points that no sample size reduces, and it is not opinion moving inside the pairing window — polls fielded on the same day show 1.35 points of it — nor a shared weekly fashion, since the within-date variance across 107 multi-poll dates is 5.99 pts².
- House effects are dispersed by 2.57 points across firms and do not average away within a firm.
The house-mix term dominates the environment's variance in every window measured. The implication for
the weights — that a 108,206-respondent panel is worth 4.1 times a 1,000-person poll rather than 108
times — was registered as c11_inverse_variance_weights and is applied as of 2026-09-27
(house_prior/0.9.0). Promoted on its pre-registered rule: a firm-clustered 95% interval of
[+0.063, +0.329] pts² on 3,765 held-out polls over 127 firms, excluding zero, reproducing the
[+0.059, +0.304] measured at registration. The metric is leave-one-poll-out against other firms'
polls in a 21-day window, so it reads no election outcome and consumed no holdout under ADR-0010.
experiments/preregistration_inverse_variance.md has the decision; the environment moved by about a
third of a House seat, which is what the registration predicted and is not the argument for it.
11.2 What is shared across races, and where it comes from
For a chamber or the electoral college, the question "which part of a race's uncertainty moves with
every other race" decides the answer, and the budget above already names it: σ²_syst is
industry-wide polling error, one number per cycle, identical across races by construction.
Everything else in a race's budget is that race's own.
So the electoral-college simulation draws σ_syst once per simulated election from the stored
budget of the races it is combining, and draws every other term per state. It does not draw the
election-to-election national swing on top of states whose margins have already been measured from
polls — which both earlier implementations did, with the effect that a 2024 forecast came out
P(D wins) = 0.585 with a 90% interval 140 electors wide in a year decided by a point and a half
(FM-66). Unpolled states take the prior shifted by the swing the polled states imply, plus the
idiosyncratic part of a state's swing and the standard error of that swing estimate.
12. Simulation
- Draw
Sposterior samples of all parameters (S set by MCSE targets, §15.3, not by round numbers). - For each draw, draw the correlated error vector
e(§10) and, for forecasts, the drift. - Transform from ALR space back to shares; compute margins; determine per-race winners.
- Aggregate: House seats, Senate seats (including holdover seats, independents' caucusing, and VP tie-break), Governors, Electoral College, national popular vote.
- Persist the full joint draw matrix, not just summaries — chamber-control probability is a functional of the joint distribution and cannot be recovered from marginals.
Convergence is checked by Monte Carlo standard error on the reported quantities, and the simulation seed is recorded (N-01).
13. Baselines the model must beat
Every sophisticated component is evaluated against, at minimum:
| Baseline | Description |
|---|---|
| B0 | Previous election result in the same geography |
| B1 | Geography partisan baseline (§8), uniform national swing applied |
| B2 | Unweighted mean of last 21 days of polls |
| B3 | Exponentially recency-weighted poll average, decay fitted by grid search |
| B4 | B3 + sample-size weighting + house-effect correction |
| B5 | Fundamentals-only (§7) |
| B6 | Simple 50/50 blend of B4 and B5 |
Plus an uncertainty baseline: B2 with an empirical error distribution taken straight from the historical distribution of (poll average − result). This is a surprisingly strong probabilistic baseline and beating it is the real test.
If the hierarchical Bayesian model cannot beat B4/B6 out of sample on log score and calibration, the simple model ships and the complex one goes back to the lab. This is a rule, not a sentiment.
14. Model uncertainty
Parameter uncertainty is not the same as being right about the model. We carry σ²_model by:
- Fitting a small set of structurally distinct specifications (different decay families, different correlation factor sets, different observation distributions).
- Evaluating each out of sample.
- Combining by stacking weights fitted on held-out log score — not by equal weights and not by in-sample marginal likelihood (which rewards overfitting under misspecification).
- Reporting the between-specification spread as an explicit variance component.
Stacking must itself be validated: an ensemble is a hypothesis too (build spec 27).
15. Estimation and diagnostics
15.1 Engine
CPU NUTS. Primary candidate: PyMC with the nutpie sampler. Fallback: CmdStan via cmdstanpy.
Selection by benchmark on the Phase 1 race (ADR-0004): ESS per wall-clock second, divergence
behaviour, and memory footprint under 16 threads. No framework is adopted for fashion.
15.2 Priors
Every prior is documented with its justification in a table in this file as it is implemented, and every prior is subjected to a prior predictive check: simulate from the prior alone and confirm the implied elections are politically plausible (no 97%-of-vote Senate winners, no 40-point national swings in a week).
15.3 Diagnostic gates (a fit that fails these is not used)
The gate inspects every sampled parameter, not a chosen few. An early implementation
summarised four hyperparameters and reported max_rhat = 1.0017 while the sampler itself was
warning that R-hat exceeded 1.01 elsewhere. That is worse than a weaker gate: it produces a
single number that reads like a guarantee over the whole model. The diagnostics now record
which parameter is worst, so a failure names its cause.
-
R̂ < 1.01on all monitored parameters -
Bulk and tail ESS > 400 per parameter
-
Zero post-warmup divergences (or a documented, investigated exception)
Divergences are treated as a modelling defect, not a sampler setting. The first fit of the Phase 1 model produced 17 of them because each pollster had an independent
HalfNormalexcess-variance parameter — unidentified for a firm with one poll, which is the classic funnel. Raisingtarget_acceptwould have quieted the symptom while leaving the variance components wrong, and those are precisely the parameters this model exists to estimate (FM-29). The fix was the hierarchical non-centred parameterisation this document already specified in §6:log τ_p ~ N(μ_τ, σ_τ²). Divergences fell to zero on most seeds. -
E-BFMI within acceptable range
-
Posterior predictive checks on held-out polls
-
Simulation-based calibration (SBC) on the synthetic DGP (TEST_PLAN.md)
16. Known weaknesses, stated up front
- Few effective observations of systematic polling error (§5). Roughly 10–20 cycle-office
units.
σ_natis weakly identified and any claim to know it precisely is false. - Non-stationarity. Polling methodology changed radically between 2000 and 2026 (landline → online panel). Historical error estimates may not transfer. Mitigation: concept-drift weighting, tested out of sample (build spec 82) — not a solution, a partial mitigation.
- House data sparsity. Most of the 435 districts are never polled. Their forecasts are fundamentals-plus-national-swing, and their uncertainty must reflect that. The chamber distribution is therefore driven largely by the correlated error structure, which is the least well-identified part of the model.
- Candidate-level idiosyncrasy (scandals, gaffes, health events) is unmodellable prospectively and enters only as idiosyncratic variance.
- Selection in the poll set. Polls are released, not sampled. Suppressed unfavourable internal polls bias the observed set in ways no weighting fixes (FM-08).
- Backtest reuse. Every pass over the same 10 cycles burns holdout. The holdout ledger (BACKTESTING.md §6) tracks the burn rate; when it is exhausted, honest evaluation requires waiting for a real election.