Adjustments and the standard test
Every adjustment between the raw polls and a published number, the one test each must pass, the latest result, and the rule that no single race's number changes the model.
Model adjustments: the registry, the standard test, and the rules for changing them
Every published number on futureballot.com is raw polls and results passed through a chain of
transformations: house effects, a prior, a variance budget, a scale correction, a blend with the
ground-up model. Until 2026-10-05 each of those was promoted on its own test, once, in isolation,
and two of them could disagree without anyone noticing. experiments/mean_bias_correction.md
concluded that no mean polling-bias correction should be applied; the promoted scale law
(experiments/scale_pre2026.json, final = 1.070 x poll - 1.744) applies a uniform shift of about
-1.74 points toward the Republican. Both were in force.
This document is the single place where that is controlled. It has four parts:
- Policy — when a specification may change, and what a challenge to one race's number triggers.
- The standard test — the one test every adjustment faces, fixed before any result was computed.
- The registry — every adjustment between raw inputs and the published number (machine-readable
copy:
data/model/adjustments.json). - The latest audit — what the standard test said about each one.
The test is implemented in src/elab/evaluate/adjustment_audit.py and run by
scripts/audit_adjustments.py; the scheduler re-runs it weekly and after each certified election.
1. Policy
(a) A challenge to one race's number is a check, never a specification change. When anyone — a reader, a reviewer, an agent, the owner — questions a single race's forecast, the response is to verify that every registered adjustment was applied to that race correctly: the right polls read as of the right time, the right walk-forward parameters, the right orientation, each step as the registry describes it. A defect found that way is a bug and is fixed (see (c)). The finding "this race's number looks wrong" is never, by itself, a reason to add, remove or retune an adjustment: an adjustment applies to every race or to none, and its merit is decided only by (b).
(b) Specification changes only through a pre-registered test that meets the standard. Adding an
adjustment, removing one, or changing how one is computed requires a pre-registration (in
experiments/, committed before results) and a result under the standard test in section 2, run on
the whole corpus. A flag from the audit is a recommendation to the coordinator, who decides through
the existing pre-registration and gate process; the audit itself changes nothing.
(c) Specification freeze from 2026-10-13 until the 2026 results are certified. From 2026-10-13 no adjustment is added, removed or re-specified, whatever a test says, until the 2026 results used for scoring are certified. Walk-forward parameters already defined as data (for example a law re-estimated on the same rule) are not specification changes. The one exception is a fix for a verified bug — code that does not do what the registry says it does, or reads data it should not — under ADR-0017. Every such fix is recorded in section 5 of this document (date, what was wrong, how it was verified, the commit, the effect on published numbers) before it is deployed.
2. The standard test
Fixed on 2026-10-05, before any audit result was computed, and committed separately from the results. Changing anything in this section is itself a specification change under (b).
2.1 What is compared
For each adjustment, two forecasts of the same race from the same standpoint with the same random seed:
- on — the full published specification, exactly as the production path computes it
(
scripts/publish_forecasts.py->elab.backtest.pipeline.forecast_race_at->scripts/apply_statewide_prior.pyfor Senate and governor;scripts/forecast_house_ground_up.pyfor the House); - off — the same specification with that one step replaced by identity (or, for a step that substitutes one estimate for another, by what it replaced). Every other step is identical.
Where an adjustment cannot act on a race (the sparse blend on a race with enough polls to fit), the off forecast equals the on forecast. Where the off forecast cannot be produced (its fit fails the convergence gate), the race is dropped from that adjustment's comparison and counted.
2.2 Walk-forward by cycle
A forecast of cycle T uses only what was knowable before T:
- parameters fitted across cycles are taken from files fitted on cycles strictly before T
(
experiments/scale_pre<T>.json,experiments/uncertainty_terms_pre<T>.json, the sparse-poll error pooled over cycles < T, the statewide ground-up prior as its gate fitted it for T, the House model's lean, error model and national term trained on cycles < T); the harness refuses any file whose recordedused_cyclesinclude T or later; - polls are read through
ObservationRepository.for_race(as_of=standpoint), the bitemporal reader every walk-forward script uses, so nothing recorded, published or released after the standpoint is seen; results enter the fundamentals prior only if knowable at the standpoint; - the realised result is attached for scoring only and is read by no forecast step (a leakage test changes it and checks every arm is unchanged).
2.3 Corpus and standpoints
- Senate and governor: every race in the stored walk-forward record of the published
specification (
single_race_latent/0.7.0election-day backtests, cycles 2006-2024, general and special), at two standpoints: eve (election day, the stored record's own standpoint) and h28 (28 days before, close to the published forecasts' horizon in October). A race whose published path produces no forecast at a standpoint (no polls, or the fit fails its gate) is not in that standpoint's corpus. Races are scored against the certified two-way margin with the refusals ofelab.evaluate.scored(a third candidate won; the two took under 80%). - House: every contested two-party district in 2016-2024 (the cycles with walk-forward error residuals behind them), one standpoint (the national term as known before the election).
2.4 Scores
Per race: absolute margin error; log loss of the win probability (clipped to [0.001, 0.999], so one confident miss contributes at most 6.9); Brier score; whether the result fell inside the 80% and 90% intervals. Reported pooled and for every cycle separately, with the number of races per cycle, so a verdict resting on few cycles is visible as such.
2.5 The pass rule
Gain = log loss without the adjustment minus log loss with it (positive means the adjustment helps). At each standpoint where the adjustment changes at least one forecast, it passes if all three hold:
- pooled gain over all races is > 0;
- the mean gain is > 0 in a strict majority of the scored cycles (cycles in which it changes at least one forecast);
- its effect does not reverse sign in any major subgroup that spans at least 3 scored cycles. The major subgroups are: midterm vs presidential cycles; Senate vs governor; competitive (standpoint polling average within 8 points: the mean of the latest five usable polls, or the prior mean for a race without polls; for the House, the on-arm central margin) vs not. A reversal is a subgroup whose pooled gain has the opposite sign from the overall gain.
An adjustment is kept if it passes at every standpoint where it changes anything, flagged for review otherwise, and untestable if the corpus cannot test it (each such case is listed with the reason). MAE, Brier and coverage are reported beside the verdict and do not decide it.
For an entry that records a choice rather than a correction (sponsored polls kept unadjusted), the comparison is the published choice (on) against the alternative named in its registry entry (off), under the same rule.
2.5a Amendment 1: spread-only terms (2026-10-05, after results)
Disclosure. This amendment was adopted on 2026-10-05 after the first audit's results were seen, at the project owner's explicit direction. It is not part of the rule committed in 7f7741d.
Why. The rule in 2.5 judges every adjustment on log loss. For a term that changes only a forecast's spread (never its centre), log loss rewards narrower forecasts whenever the scored cycles hold few upsets, so it would remove terms whose intervals are correctly calibrated. ADR-0016 already requires spread changes to be judged on coverage (and CRPS); the standard as first written contradicted it.
Rule. For the terms in SPREAD_ONLY (src/elab/evaluate/adjustment_audit.py: design effect,
measured sparse-poll error, sigma_nat, sigma_race, drift to election day, and the House national,
state and race-type error terms), at each standpoint where the term changes a forecast, it passes if
(1) with it, pooled 80% and 90% interval coverage are each within 0.03 of nominal, and (2) switching
it off does not bring total coverage error (|cov80 - 0.80| + |cov90 - 0.90|) closer to nominal by
more than 0.01. Per-cycle majority and subgroup tests are not applied to these terms: with 6 to 68
races a cycle, a single cycle's coverage has a standard error of 4 to 12 points. CRPS is not used
because the stored summaries hold quantiles, not full distributions; adding it is listed as work.
The log-loss verdict is still computed and stored beside it.
Result (re-scored from the stored audit, no new fits;
experiments/spread_amendment1_rescore_2026-10-05.json):
| Term | Verdict | 80%/90% coverage with it | without it |
|---|---|---|---|
| sigma_nat | keep | 0.793/0.905 (eve), 0.816/0.900 (4 weeks) | 0.769/0.891, 0.791/0.894 |
| sigma_race | keep | same | 0.491/0.607, 0.674/0.802 |
| drift to election day | keep | 0.816/0.900 (4 weeks) | 0.747/0.864 |
| design effect | keep | 0.790/0.903, 0.815/0.898 | 0.790/0.906, 0.812/0.901 |
| measured sparse-poll error | keep | 0.793/0.905, 0.816/0.900 | 0.775/0.891, 0.805/0.889 |
| House national, state, race-type error | flag | 0.938/0.974 (eve) | 0.92-0.93/0.96-0.97 |
The House terms fail condition (1): House intervals are too wide (90% intervals hold 97.4% of outcomes). Removing any one term does not fix that, so the remedy is a recalibration of the district error model under its own pre-registration, not deleting a term.
2.6 The contradiction check
Each registry entry lists the claims its evidence makes (claims in data/model/adjustments.json):
a quantity and a conclusion (correct or do_not_correct). Any two claims about the same quantity
with opposite conclusions are reported as a contradiction, whatever the status of the entries
involved (an active adjustment against a rejected experiment is exactly the case that went
unnoticed). A contradiction is resolved only by a pre-registered test under (b), or by recording in
the registry why the two claims are about different quantities.
2.7 Effect on today's forecasts
For each adjustment: the median and maximum absolute change in the election-day margin and in the win probability across today's published races when that one step is switched off, computed in memory from the same inputs the published run used. Nothing is persisted.
2.8 Automation and alerts
scripts/audit_adjustments.py --if-due --alert runs daily from the scheduler and does the work when
seven days have passed since the last audit or when the set of certified results has changed (that
is, after each certified election). It sends one phone alert (ntfy, as the integrity job does) when
an adjustment that was kept is now flagged, or when a contradiction appears that the previous audit
did not report. Fits are cached by their exact inputs outside the repository, so a re-run refits
only races whose polls, results or parameters changed.
3. The registry
One entry per transformation, found by reading the code paths that produce the published forecasts: scripts/scheduler.py jobs forecast-senate/forecast-governor (scripts/publish_forecasts.py -> src/elab/backtest/pipeline.py -> src/elab/models/*, src/elab/simulate/forecast.py), apply-statewide-prior (scripts/apply_statewide_prior.py), the chamber jobs (scripts/publish_chamber.py) and chamber-house (scripts/forecast_house_ground_up.py and the national term from scripts/backtest_national.py). The machine-readable copy, with parameters, evidence files and the claims the contradiction check reads, is data/model/adjustments.json.
Today's effect is measured on the published forecasts of 2026-10-05 (59 Senate and governor races with a poll-based run; the 12 published from the ground-up prior alone are affected only by that prior; 418 contested House districts) by switching the one step off in memory: the absolute change in the election-day margin (points) and in the win probability, median and maximum across races, with the race where the margin moved most. Nothing was persisted.
Senate and governor
poll_selection_asof — Which polls a forecast reads (active)
- What it does: A forecast reads only polls recorded, published and (for reconstructed polls) released by its standpoint, the headline question only, not withdrawn or superseded; a poll with several versions keeps the first one stored.
- Where:
src/elab/repo/observations.py:ObservationRepository.for_race - Parameters: ELAB_RECON_RELEASE_LAG_DAYS=0, ELAB_RECON_RELEASE_BASIS=field_end
- Evidence: docs/BACKTESTING.md; tests/leakage/test_asof_canary.py; tests/leakage/test_release_basis.py
- Last tested before this audit: 2026-10-05. Protocol: Leakage tests (canary poll after the as_of must change nothing); a data rule, not a model adjustment.
- Today's effect: -
- Standard test (2026-10-05): UNTESTABLE — A knowability rule, not an adjustment: switching it off is leakage, not an alternative forecast. Guarded by the leakage tests.
lv_rv_handling — Likely- vs registered-voter polls (active)
- What it does: None. Likely-voter, registered-voter and adult samples enter alike; the population column is stored and read by nothing in the model. When a release is listed in several populations, whichever version was stored first is the one used (the duplicate rule keys on pollster and fieldwork). A firm's typical lean, including any from its population, is absorbed by its house effect.
- Where:
src/elab/repo/observations.py (population read, unused);src/elab/ingest/wikipedia_discovery.py:_already_held - Parameters: none
- Evidence: docs/STATISTICAL_MODEL.md section 4 (describes an LV/RV term g_pop that is not implemented); docs/FAILURE_MODES.md FM-79 (the sigma_lv placeholder retired into sigma_race)
- Last tested before this audit: never. Protocol: Never tested.
- Today's effect: -
- Standard test (2026-10-05): UNTESTABLE — Nothing is applied, so there is nothing to switch off; 7,242 of 11,258 historical statewide polls record no population at all, so an LV/RV term cannot be fitted walk-forward from this corpus. Adding one would be a new specification under policy (b).
sponsor_handling — Campaign- and party-sponsored polls (active)
- What it does: Sponsored polls (a campaign, party committee, PAC or interest group) are kept and treated like any other poll; the sponsor's lean is left to the pollster's house effect. Roughly half of live statewide polls are flagged sponsored (327 of 676 when counted on 2026-10-05). Tested against the alternative of dropping them.
- Where:
src/elab/backtest/pipeline.py:forecast_race_at (exclude_partisan=False for statewide);src/elab/ingest/poll_sponsors.py - Parameters: exclude_partisan=False
- Evidence: experiments/preregistration_polled_nonpartisan.md; experiments/house_polled_nonpartisan.md
- Last tested before this audit: 2026-09-30. Protocol: House districts only: seat MAE 2018-2024 against the champion; excluding sponsored polls was refused (8.1 vs 6.0). Never tested statewide.
- Today's effect: margin median 0.97 / max 27.53 (US-SD); win prob median 0.021 / max 0.399; 50 of 59 races changed
- Standard test (2026-10-05): FLAG
pollster_house_effects — Pollster house effects (sum-to-zero) (active)
- What it does: Each pollster in a race gets its own lean, estimated from that race's polls; the leans are constrained to average zero across the race's polls (weighted by how many each firm released), so the latent path tracks 'the average pollster'.
- Where:
src/elab/models/latent/single_race.py:build_model (house, house_sd, sum-to-zero by volume) - Parameters: house_sd ~ HalfNormal(2.5); weights = poll counts
- Evidence: docs/STATISTICAL_MODEL.md section 5; docs/FAILURE_MODES.md FM-13; experiments/house_effects.md; experiments/batch03_pooled_house.md
- Last tested before this audit: 2026-09-24. Protocol: Never gated itself; the pooled cross-race alternative was tested (log loss, 11 cycles) and rejected (+0.0013, 5 of 11 cycles).
- Today's effect: margin median 0.04 / max 0.84 (US-ID); win prob median 0.001 / max 0.036; 43 of 59 races changed
- Standard test (2026-10-05): FLAG
design_effect — Design effect 1.6 on poll sampling error (active)
- What it does: Every poll's sampling variance is multiplied by 1.6 before it enters the fit, on the view that reported margins of error understate real error.
- Where:
src/elab/models/latent/single_race.py:sampling_sd_points;src/elab/models/sparse.py:DESIGN_EFFECT - Parameters: 1.6 (constant)
- Evidence: docs/STATISTICAL_MODEL.md section 4.1; experiments/poll_dispersion*.json
- Last tested before this audit: never. Protocol: Never tested; a hard-coded constant.
- Today's effect: margin median 0.06 / max 0.73 (US-SD); win prob median 0.001 / max 0.021; 43 of 59 races changed
- Standard test (2026-10-05): FLAG
fundamentals_prior — Fundamentals prior on the race intercept (active)
- What it does: The latent path's starting level gets an informative prior from the state's previous results (same seat where possible) with a width from how much such margins move, instead of a vague Normal(0, 15) centred on a tie.
- Where:
scripts/publish_forecasts.py (prior_for, measure_volatility);src/elab/models/fundamentals/prior.py:prior_for;src/elab/models/latent/single_race.py:build_model start_margin - Parameters: volatility from cycles before the target
- Evidence: docs/DECISIONS.md ADR-0007; docs/FAILURE_MODES.md FM-84
- Last tested before this audit: 2026-09-25. Protocol: Shipped as a fix (ADR-0017), not gated; 284 races: CRPS 3.928 -> 3.871, MAE 5.46 -> 5.40, better in 5 of 5 cycles.
- Today's effect: margin median 0.10 / max 6.94 (US-LA); win prob median 0.002 / max 0.082; 57 of 59 races changed
- Standard test (2026-10-05): KEEP
min_polls_rule — Four polls to fit a trajectory (active)
- What it does: A race with four or more usable polls is fitted by the latent model; one with one to three goes to the sparse blend; one with none is not a poll forecast.
- Where:
scripts/publish_forecasts.py --min-polls 4;src/elab/backtest/pipeline.py:forecast_race_at - Parameters: min_polls = 4 (the remote worker fits from 3)
- Evidence: experiments/preregistration_min_polls.md; experiments/prior_only_fallback_ignores_sparse_polls.md
- Last tested before this audit: 2026-09-23. Protocol: k=3 (c6) tested: passed (log loss -0.0395) and superseded by c7; the value 4 itself was never tested.
- Today's effect: -
- Standard test (2026-10-05): UNTESTABLE — A threshold with no 'off' state: below three polls the latent model cannot be fitted at all. Comparing 3 against 4 is a specification change under policy (b).
sparse_blend — Sparse-race blend (c7) (active)
- What it does: A race with one to three polls is forecast by combining the fundamentals prior with the average of its polls, each weighted by its precision, instead of the prior alone.
- Where:
src/elab/models/sparse.py:sparse_estimate;src/elab/backtest/pipeline.py:forecast_race_at (len(usable) < min_polls) - Parameters: prior (mean, sd); poll sd as sparse_measured_sd
- Evidence: experiments/preregistration_sparse_blend.md; experiments/prior_only_fallback_ignores_sparse_polls.md
- Last tested before this audit: 2026-09-23. Protocol: Pre-registered, confirmation cycles 2018-2024, log loss, BH q <= 0.10: -0.0474, q=0.0002, 4 of 4 cycles. Promoted.
- Today's effect: margin median 0.00 / max 11.78 (US-LA); win prob median 0.000 / max 0.172; 15 of 59 races changed
- Standard test (2026-10-05): KEEP
sparse_measured_sd — Realised error of a sparse race's poll average (FM-104) (active)
- What it does: In the sparse blend a handful of polls is weighted by how far such averages have actually landed from results in earlier cycles (about 9-10 points), not by their sampling error (about 4).
- Where:
src/elab/models/sparse.py:pooled_realised_sd;scripts/publish_forecasts.py (sparse_poll_variance.json) - Parameters: experiments/sparse_poll_variance.json, pooled over cycles before the target
- Evidence: experiments/sparse_blend_poll_weight.md; docs/FAILURE_MODES.md FM-104; experiments/sparse_variance.md
- Last tested before this audit: 2026-09-28. Protocol: Shipped as a defect fix: 67 sparse races, MAE 8.01 -> 6.34, 90% coverage 0.791 -> 0.925. As challenger c9 (sparse_variance.md, 2026-09-25) the same change was refused on log loss (+0.0150, 0 of 3 folds).
- Today's effect: margin median 0.00 / max 7.44 (US-LA); win prob median 0.000 / max 0.105; 15 of 59 races changed
- Standard test (2026-10-05): FLAG
sigma_nat — Shared national polling error (sigma_nat) (active)
- What it does: Every race's election-day spread includes an error shared by all races in a cycle, the size of past cycles' average polling misses (2.44 points for 2026).
- Where:
src/elab/simulate/forecast.py:election_day;scripts/publish_forecasts.py (terms['systematic']['sigma_nat']) - Parameters: experiments/uncertainty_terms_pre<cycle>.json
- Evidence: experiments/uncertainty_terms.md; docs/FAILURE_MODES.md FM-82
- Last tested before this audit: 2026-09-23. Protocol: Estimated (variance components over earlier cycles), never gated.
- Today's effect: margin median 0.08 / max 0.46 (US-OR); win prob median 0.005 / max 0.020; 59 of 59 races changed
- Standard test (2026-10-05): FLAG
sigma_race — Race-specific residual error (sigma_race) (active)
- What it does: Every race's spread includes a race-specific error beyond its polls (4.97 points for 2026), the residual of the same variance-components fit.
- Where:
src/elab/simulate/forecast.py:election_day;scripts/publish_forecasts.py (terms['systematic']['sigma_race']) - Parameters: experiments/uncertainty_terms_pre<cycle>.json
- Evidence: experiments/uncertainty_terms.md; docs/FAILURE_MODES.md FM-79; docs/FAILURE_MODES.md FM-89
- Last tested before this audit: 2026-09-26. Protocol: Restored as a fix (FM-79): 247 races, 90% coverage 0.680 -> 0.895 walk-forward. Not gated.
- Today's effect: margin median 0.36 / max 2.04 (US-OR); win prob median 0.023 / max 0.106; 59 of 59 races changed
- Standard test (2026-10-05): FLAG
drift_law — Drift to election day (active)
- What it does: A forecast made before election day is widened by how far races have historically moved between that horizon and the vote: variance ah^b, h in days (1.477h^0.691 for 2026, about 7.6 points of sd at four weeks).
- Where:
scripts/publish_forecasts.py (drift VarianceTerm);src/elab/simulate/forecast.py:election_day - Parameters: experiments/uncertainty_terms_pre<cycle>.json drift a, b
- Evidence: experiments/uncertainty_terms.md
- Last tested before this audit: 2026-09-23. Protocol: Estimated from historical trajectories, never gated.
- Today's effect: margin median 0.21 / max 1.22 (US-OR); win prob median 0.014 / max 0.057; 59 of 59 races changed
- Standard test (2026-10-05): FLAG
student_t_shocks — Fat-tailed (Student-t, nu=5) election-day errors (active)
- What it does: The non-poll error terms are drawn from a Student-t with 5 degrees of freedom rather than a normal, so large misses are more likely than a bell curve allows.
- Where:
src/elab/simulate/forecast.py:_simulate (nu=5.0) - Parameters: nu = 5 (constant)
- Evidence: docs/STATISTICAL_MODEL.md section 10; docs/FAILURE_MODES.md FM-12
- Last tested before this audit: never. Protocol: Never tested; a constant.
- Today's effect: margin median 0.05 / max 0.28 (US-AK); win prob median 0.003 / max 0.023; 59 of 59 races changed
- Standard test (2026-10-05): KEEP
scale_law — Scale law c15 (slope and intercept together) (active)
- What it does: The whole forecast distribution is translated so its centre becomes a + b x centre, with a and b fitted on earlier cycles' final polling averages against results (2026: -1.744 + 1.070 x centre). Reported for reference; the two parts are judged separately below.
- Where:
src/elab/models/calibration/scale.py:ScaleLaw.apply;src/elab/backtest/pipeline.py:forecast_race_at - Parameters: experiments/scale_pre<cycle>.json
- Evidence: experiments/preregistration_scale.md; experiments/scale_compression.md; experiments/scale_calibration_2016.md; experiments/scale_intercept_check_2026-10-05.md
- Last tested before this audit: 2026-09-26. Protocol: Pre-registered, walk-forward 2016-2024, CRPS: -0.530 against threshold 0.073, 4 of 5 folds, log loss -0.0088. Promoted as 0.6.0.
- Today's effect: margin median 1.31 / max 2.85 (US-MT); win prob median 0.028 / max 0.115; 59 of 59 races changed
- Standard test (2026-10-05): FLAG
scale_slope — Scale law: slope (anti-compression) (active)
- What it does: Stretches every forecast away from a tie by the fitted slope (1.070 for 2026), because final polling averages have been more compressed than results.
- Where:
src/elab/models/calibration/scale.py:ScaleLaw.apply - Parameters: experiments/scale_pre<cycle>.json slope
- Evidence: experiments/preregistration_scale.md; experiments/scale_compression.md; experiments/compression_mechanism.md
- Last tested before this audit: 2026-09-26. Protocol: Tested only jointly with the intercept (c15).
- Today's effect: margin median 0.43 / max 1.26 (US-CA); win prob median 0.007 / max 0.019; 59 of 59 races changed
- Standard test (2026-10-05): KEEP
scale_intercept — Scale law: intercept (a uniform shift) (active)
- What it does: Moves every forecast by the fitted intercept (-1.744 points, toward the Republican, for 2026). This is a mean polling-bias correction.
- Where:
src/elab/models/calibration/scale.py:ScaleLaw.apply - Parameters: experiments/scale_pre<cycle>.json intercept (walk-forward values from +1.52 to -1.96)
- Evidence: experiments/preregistration_scale.md; experiments/scale_intercept_check_2026-10-05.md
- Last tested before this audit: 2026-10-05. Protocol: Never tested on its own before; entry 1 of this audit (scale_intercept_check_2026-10-05) found it helps CRPS overall but not log loss, helps in presidential years and hurts in midterms.
- Today's effect: margin median 1.37 / max 1.74 (US-ID); win prob median 0.028 / max 0.113; 59 of 59 races changed
- Standard test (2026-10-05): FLAG
- Challenged 2026-10-05 (pre-registered at b5a423d;
experiments/mean_bias_challengers_2026-10-05.md): slope only (S1) and a same-kind intercept (S2) both not promoted on the citable folds 2006-2016. The champion's intercept stays.
mean_bias_correction — Mean polling-bias correction (rejected) (rejected)
- What it does: Subtracting the average past polling miss from every forecast. Rejected; listed for its evidence.
- Where: -
- Parameters: n/a
- Evidence: experiments/mean_bias_correction.md; experiments/bias_centre.md; experiments/preregistration_bias_centre.md; experiments/preregistration_recency_centre.md
- Last tested before this audit: 2026-09-29. Protocol: mean_bias_correction.md (2026-09-23): out-of-sample |error| 3.22 -> 3.09, judged noise. c8_bias_centre (2026-09-25): rejected three times on log loss. c19 recency centre (2026-09-29): not promoted.
- Today's effect: not in the published path
- Standard test (2026-10-05): UNTESTABLE — Not in the published path (rejected); its question is answered by scale_intercept.
statewide_groundup_blend — Statewide ground-up prior blend (0.8.1) (active)
- What it does: The election-day centre of every Democrat-vs-other race is averaged with a ground-up prediction (a gradient-boosted model of the state's lean from its results, demographics, economy, candidates and money, plus the national House environment N3), weighted by each one's variance; the spread is kept and the win probability recomputed.
- Where:
scripts/apply_statewide_prior.py:apply;scripts/gate_statewide.py:prior_for_cycle - Parameters: data/national/state_panel.json, data/national/backtest.json (N3), prior variance from walk-forward residuals
- Evidence: experiments/statewide_gate_preregistration.md; experiments/statewide_gate_result.md
- Last tested before this audit: 2026-10-01. Protocol: Pre-registered gate, selection 2016-2022, confirmation 2024; Brier and MAE (not log loss): Brier 0.0677 -> 0.0585, MAE 5.60 -> 5.38 on 175 races; most of the gain is 2020.
- Today's effect: margin median 1.46 / max 6.21 (US-ID); win prob median 0.019 / max 0.305; 55 of 59 races changed
- Standard test (2026-10-05): FLAG
prior_only_publication — Unpolled races published from the ground-up prior (FM-140) (active)
- What it does: A Senate or governor race with no poll at all is published from the ground-up prior alone, with the prior's full measured error as its spread, labelled 'fundamentals only, no polls'.
- Where:
scripts/apply_statewide_prior.py:publish_prior_only - Parameters: prior sd from walk-forward residuals
- Evidence: docs/FAILURE_MODES.md FM-140; experiments/unpolled_seat_tails.md
- Last tested before this audit: 2026-10-02. Protocol: No gate; justified by the prior's own error (MAE 8.3 on 2016-2022, 6.7 in 2024) and a tail check (0 of 87 seats favoured by 20+ flipped).
- Today's effect: -
- Standard test (2026-10-05): UNTESTABLE — Switching it off leaves the race unpublished, not forecast differently: there is no 'off' forecast to score. The prior's own accuracy is the evidence.
shock_rules — Shock uncertainty and centre-shift rules (off) (inactive)
- What it does: Widening or shifting a race after a confirmed severe event (a withdrawal, charge, scandal). Built, registered, failed their tests, switched off.
- Where:
scripts/apply_statewide_prior.py (ELAB_SHOCK_UNCERTAINTY_SD, ELAB_SHOCK_SHIFT_POINTS unset);src/elab/models/shock.py - Parameters: unset
- Evidence: experiments/shock_uncertainty_result.md; experiments/shock_shift_result.md
- Last tested before this audit: 2026-10-04. Protocol: Pre-registered; both failed.
- Today's effect: not in the published path
- Standard test (2026-10-05): UNTESTABLE — Not in the published path (switched off after failing their own pre-registered tests).
Chambers
chamber_correlation — Chamber simulation: correlated errors and bias sensitivity (active)
- What it does: Senate and governor control probabilities draw one shared national error (and regional and office terms) per simulated election across all races; the model's measured lean is shown only as a sensitivity beside the headline, never applied to it.
- Where:
scripts/publish_chamber.py;src/elab/models/correlation/structure.py:estimate_structure;src/elab/simulate/chamber.py - Parameters: estimated from history at run time
- Evidence: experiments/chamber_gate_preregistration.md; experiments/chamber_gate_result.md
- Last tested before this audit: 2026-10-01. Protocol: Chamber gate (pre-registered).
- Today's effect: -
- Standard test (2026-10-05): UNTESTABLE — Acts on chamber totals, not race forecasts: one chamber outcome per cycle gives five to ten scored points, too few for the per-cycle rule. Race-level numbers are unaffected.
House (ground-up model)
house_national_adjustments — House: national term N3 instead of the generic ballot at face value (active)
- What it does: The House model's national environment is N3 (the generic-ballot average with its mean past miss removed, blended with a fundamentals regression) rather than the generic ballot as it stands. Both corrections switched off together.
- Where:
scripts/backtest_national.py (n1, n2, blend);scripts/forecast_house_ground_up.py (NAT[2026]['N3']) - Parameters: data/national/backtest.json; 2026: N0 9.38 -> N3 7.31
- Evidence: experiments/national_environment.md; experiments/district_gate_result.md
- Last tested before this audit: 2026-10-01. Protocol: National: walk-forward MAE 2004-2024, N0 3.01, N3 2.16. District gate tested N3 only as part of a package (lean model, national term and errors changed together).
- Today's effect: margin median 2.02 / max 2.14 (US-NY-23); win prob median 0.004 / max 0.098; 418 of 418 races changed; seats 243.9 -> 253.5, P(D majority) 0.984 -> 0.989
- Standard test (2026-10-05): FLAG
house_gb_bias_correction — House: generic-ballot bias correction inside N3 (active)
- What it does: Subtracts the generic ballot's average past miss (2.62 points toward the Democrat, from 1998-2024) before it enters the national term.
- Where:
scripts/backtest_national.py:n2 - Parameters: mean(generic ballot - result) over 1998..t-2
- Evidence: experiments/national_environment.md; experiments/house_generic_ballot_bias.md; experiments/preregistration_generic_ballot_bias.md
- Last tested before this audit: 2026-10-01. Protocol: Never isolated: national_environment.md (N2 MAE 2.44 vs 3.01) ran the same day house_generic_ballot_bias.md refused the same correction for the old House model (environment MAE 2.72 vs 2.73).
- Today's effect: margin median 1.69 / max 1.82 (US-NY-23); win prob median 0.003 / max 0.083; 418 of 418 races changed; seats 243.9 -> 251.8, P(D majority) 0.984 -> 0.994
- Standard test (2026-10-05): FLAG
- Challenged 2026-10-05 (pre-registered at b5a423d;
experiments/mean_bias_challengers_2026-10-05.md): no correction (H1) and a same-kind correction (H2) not promoted. The ledger leaves only 2016 citable (2018-2024 are burned), so neither could be promoted, and on 2016 both have worse seat error. The midterm-only generic-ballot miss (2.70, seven midterms) is the same size as the pooled one (2.62), so a same-kind correction leaves 2026 where it is.
house_fundamentals_in_national — House: fundamentals regression blended into the national term (active)
- What it does: Blends the bias-corrected generic ballot with a regression on midterm penalty, approval, income growth and the previous House vote, weighted by each one's past error.
- Where:
scripts/backtest_national.py:n1 and the N3 blend - Parameters: N1 coefficients refitted on 1960..t-2
- Evidence: experiments/national_environment.md
- Last tested before this audit: 2026-10-01. Protocol: Walk-forward national MAE 2004-2024 (N2 2.44, N3 2.16); not isolated at the district level.
- Today's effect: margin median 0.56 / max 0.63 (US-NY-2); win prob median 0.001 / max 0.028; 418 of 418 races changed; seats 243.9 -> 241.4, P(D majority) 0.984 -> 0.972
- Standard test (2026-10-05): FLAG
house_nat_sd — House: national uncertainty (active)
- What it does: Every simulated House election draws one national environment around N3 with the national estimator's walk-forward error (2.80 points for 2026).
- Where:
scripts/gate_district_model.py:nat_sd;src/elab/models/district/model.py:simulate - Parameters: RMSE of N3's errors before the cycle
- Evidence: experiments/district_gate_preregistration.md
- Last tested before this audit: 2026-10-01. Protocol: Part of the district gate package; not isolated.
- Today's effect: margin median 0.08 / max 0.57 (US-CA-30); win prob median 0.002 / max 0.020; 418 of 418 races changed; seats 243.9 -> 243.1, P(D majority) 0.984 -> 1.000
- Standard test (2026-10-05): FLAG
house_state_sd — House: shared state error (active)
- What it does: Districts in the same state share an error (3.55 points for 2026), measured from walk-forward residuals.
- Where:
src/elab/models/district/model.py:ErrorModel.from_residuals - Parameters: walk-forward residuals 2014..t-2
- Evidence: experiments/district_gate_preregistration.md
- Last tested before this audit: 2026-10-01. Protocol: Part of the district gate package; not isolated.
- Today's effect: margin median 0.01 / max 0.06 (US-MO-3); win prob median 0.003 / max 0.024; 418 of 418 races changed; seats 243.9 -> 243.0, P(D majority) 0.984 -> 0.985
- Standard test (2026-10-05): FLAG
house_idio_buckets — House: district error by race type (active)
- What it does: A district's own error is sized by its type: incumbent on the same lines (7.25), open seat (7.86) or new lines (8.11), rather than one size for all.
- Where:
src/elab/models/district/model.py:bucket, ErrorModel - Parameters: walk-forward residuals 2014..t-2
- Evidence: experiments/district_gate_preregistration.md
- Last tested before this audit: 2026-10-01. Protocol: Part of the district gate package; not isolated.
- Today's effect: margin median 0.00 / max 0.01 (US-CA-36); win prob median 0.001 / max 0.014; 418 of 418 races changed; seats 243.9 -> 243.5, P(D majority) 0.984 -> 0.983
- Standard test (2026-10-05): FLAG
house_nat_swing — House: district-specific response to the national swing (active)
- What it does: The lean model takes the national swing as a feature, so each district's response to a wave is learned rather than uniform. Off = uniform swing.
- Where:
scripts/backtest_district_model.py:feature_columns (nat_swing);src/elab/models/district/model.py:lean_grid - Parameters: learned
- Evidence: scripts/backtest_district_model.py (pre-registered docstring); experiments/district_gate_result.md
- Last tested before this audit: 2026-10-01. Protocol: Part of the district gate package; not isolated.
- Today's effect: margin median 0.54 / max 3.67 (US-TX-16); win prob median 0.001 / max 0.082; 418 of 418 races changed; seats 243.9 -> 243.8, P(D majority) 0.984 -> 0.981
- Standard test (2026-10-05): FLAG
house_uncontested_rule — House: seats without a two-party contest (active)
- What it does: A district without both a Democrat and a Republican on the ballot is given, unsimulated, to the one major party with a nominee on record (qual_D or qual_R present). A district where neither side, or both, has a nominee on record has an unknown field and is simulated from its lean (FM-173; before, it was an automatic Republican seat).
- Where:
src/elab/models/district/model.py:split_one_party;scripts/forecast_house_ground_up.py (fixed_d, one_party) - Parameters: none
- Evidence: scripts/forecast_house_ground_up.py
- Last tested before this audit: never. Protocol: Never tested; the gate used actual winners for these seats instead.
- Today's effect: -
- Standard test (2026-10-05): UNTESTABLE — These seats get no district forecast; the walk-forward gate assigned them from actual winners (an oracle), so there is no historical 'on' arm to score. Correctness under policy (a): the old rule made a seat with no nominee on record a Republican seat (LA-2 before FM-155 repaired Louisiana's rosters); fixed by FM-173, no effect on today's House (no such seat), +0.99 seats on the pre-FM-155 data.
house_district_polls — House: district poll update (off) (inactive)
- What it does: Shifting a polled district's lean by its polls. Off by default after its adversarial review.
- Where:
scripts/forecast_house_ground_up.py --polls;src/elab/models/district/polls.py - Parameters: off
- Evidence: experiments/district_polls_review.md
- Last tested before this audit: 2026-10-01. Protocol: Rejected by adversarial review.
- Today's effect: not in the published path
- Standard test (2026-10-05): UNTESTABLE — Not in the published path.
4. The latest audit (2026-10-05)
Full report: experiments/adjustment_audit_2026-10-05.md (per-cycle tables for every adjustment at
every standpoint, subgroups, races dropped and why) and .json. The run's header names commit
7f7741d (the pre-registration) because the harness was not yet committed; the harness is commit
8e93e76, identical in its arithmetic to what ran (that commit adds only quieter worker logs and
--check-race). Thirty minutes on Skynet-Three:
3,860 walk-forward fits in a process pool, nothing persisted. The corpus is 459 Senate and governor
races from 2006-2024 at each standpoint (377 scoreable on the eve, 359 four weeks out; 78 are out of
scoring because elab.evaluate.scored refuses them, 77 of those for an uncertified result) and
1,971 contested House districts from 2016-2024.
The engine reproduces the published numbers. Its "on" arm matches today's stored Senate and governor forecasts to a median 0.002 points (98% within 0.1; largest gap 0.51) and today's House file exactly (418 districts, 243.9 seats, P(D majority) 0.984). Against the stored backtests it matches to a median 0.03 points; of the 69 races more than 0.1 apart, 42 (and all of the largest, up to 4.1) are thin sparse-blend races whose stored backtests predate the 2026-09-29 refit of the sparse-poll error and the poll backfills since. The audit scores today's specification on today's data, as it should.
Verdicts
| adjustment | verdict | eve: gain, cycles improved | four weeks out: gain, cycles improved | today: margin shift median / max | today: win-prob shift median / max |
|---|---|---|---|---|---|
| fundamentals_prior | keep | +0.0015, 8/10 | +0.0026, 8/10 | 0.10 / 6.94 | 0.002 / 0.082 |
| sparse_blend | keep | +0.0334, 8/8 | +0.0568, 10/10 | 0.00 / 11.78 | 0.000 / 0.172 |
| student_t_shocks | keep | +0.0022, 8/10 | +0.0030, 8/10 | 0.05 / 0.28 | 0.003 / 0.023 |
| scale_slope | keep | +0.0054, 8/10 | +0.0068, 7/10 | 0.43 / 1.26 | 0.007 / 0.019 |
| scale_intercept | flag | -0.0065, 6/10; reverses (presidential) | -0.0025, 6/10; reverses (presidential, Senate) | 1.37 / 1.74 | 0.028 / 0.113 |
| scale_law (both parts) | flag | -0.0016, 6/10; reverses | +0.0037, 6/10; reverses (midterm) | 1.31 / 2.85 | 0.028 / 0.115 |
| statewide_groundup_blend | flag | +0.0078, 5/5 (passes) | +0.0108, 3/5; reverses (governor, lopsided) | 1.46 / 6.21 | 0.019 / 0.305 |
| sponsor_handling (keep sponsored polls) | flag | -0.0020, 6/8; reverses | +0.0040, 6/8; reverses (midterm) | 0.97 / 27.53 | 0.021 / 0.399 |
| pollster_house_effects | flag | +0.0016, 4/10 | +0.0001, 5/10; reverses | 0.04 / 0.84 | 0.001 / 0.036 |
| design_effect | flag | -0.0009, 3/10 | -0.0023, 3/10 | 0.06 / 0.73 | 0.001 / 0.021 |
| sparse_measured_sd | flag | -0.0011, 0/3 | -0.0014, 0/3 | 0.00 / 7.44 | 0.000 / 0.105 |
| sigma_nat | flag | -0.0025, 2/10 | -0.0027, 3/10 | 0.08 / 0.46 | 0.005 / 0.020 |
| sigma_race | flag | +0.0137, 5/10; reverses | -0.0095, 3/10; reverses | 0.36 / 2.04 | 0.023 / 0.106 |
| drift_law | flag | (no effect on the eve) | -0.0080, 3/10 | 0.21 / 1.22 | 0.014 / 0.057 |
| house_national_adjustments (N3 vs face value) | flag | +0.0049, 2/5 | - | 2.02 / 2.14; seats 243.9 -> 253.5 | 0.004 / 0.098 |
| house_gb_bias_correction | flag | +0.0014, 2/5; reverses (close races) | - | 1.69 / 1.82; seats 243.9 -> 251.8 | 0.003 / 0.083 |
| house_fundamentals_in_national | flag | +0.0002, 2/5; reverses | - | 0.56 / 0.63; seats -> 241.4 | 0.001 / 0.028 |
| house_nat_sd | flag | -0.0020, 0/5 | - | 0.08 / 0.57 | 0.002 / 0.020 |
| house_state_sd | flag | -0.0035, 0/5 | - | 0.01 / 0.06 | 0.003 / 0.024 |
| house_idio_buckets | flag | -0.0019, 2/5 | - | 0.00 / 0.01 | 0.001 / 0.014 |
| house_nat_swing | flag | +0.0002, 2/3; reverses (close races) | - | 0.54 / 3.67 | 0.001 / 0.082 |
Gain is log loss without minus with, per race (positive: the adjustment helps). Four of 21 tested adjustments pass the standard. Nine entries cannot be tested (section 3 gives each reason).
What the flags mean, and the recommendation for each
These are recommendations to the coordinator. Nothing has been changed; any change goes through a pre-registration under policy (b), and from 2026-10-13 the freeze in (c) applies.
- scale_intercept: fails, and the failure is the mean-bias contradiction. Pooled log loss is
worse with it at both standpoints. By kind of year it helps in presidential years (+0.0101 on the
eve, five cycles) and hurts in midterms (-0.0186, five cycles; 2014 -0.043, 2022 -0.038). 2026 is
a midterm. Today it moves every Senate and governor race 1.37 points toward the Republican
(median; 1.74 before the ground-up blend dilutes it) and win probabilities by up to 11 points.
This agrees with entry 1 (
experiments/scale_intercept_check_2026-10-05.md, CRPS protocol). Recommendation: pre-register, before the 2026-10-13 freeze, a test of the slope alone (and, as the alternative entry 1 proposes, a same-kind intercept) against the current law, decided under this standard. The slope passes and should stay whatever is decided about the intercept. - The variance terms (sigma_nat, sigma_race four weeks out, drift_law, design_effect, sparse_measured_sd, house_nat_sd, house_state_sd, house_idio_buckets) fail in the same way: per-race log loss prefers narrower distributions. Interval coverage with them is on target for Senate and governor (90% intervals cover 0.905 on the eve and 0.900 four weeks out; without the drift law 0.864, without sigma_nat 0.891), so removing them would buy log loss with overconfident intervals. The House is different: its 90% intervals cover 0.974 and its 80% intervals 0.938, so the House district error model is genuinely too wide. Recommendation: no removal. Note the limitation of the standard as registered: it judges every adjustment on log loss, while ADR-0016 says a change to a distribution's spread should be judged on CRPS and coverage; amending the standard to do that for spread-only terms is a specification change for the coordinator to decide under (b), not something to adopt after seeing these results. Separately, pre-register a recalibration test of the House district error model (state and idiosyncratic terms).
- statewide_groundup_blend: passes on the eve (5 of 5 cycles), fails four weeks out on two near-zero subgroup reversals (governor -0.0001, lopsided races -0.0023), with its gains concentrated in presidential years (2020 +0.078). Its prior's national term is N3, which carries the generic-ballot bias correction (contradictions below). Recommendation: keep; include it in the midterm-bias test of item 1, since part of what it adds is the same level correction.
- House national term. N3 beats the generic ballot at face value overall (+0.0049) but only in 2 of 5 cycles: it helps in presidential years and hurts in both midterms (2018, 2022). The generic-ballot bias correction alone is worth 7.9 seats today (243.9 with it, 251.8 without) and hurts close races. Recommendation: the same question as item 1 on the House side: pre-register a test of the bias correction in midterm years. Five House cycles (two midterms) is thin; say so.
- sponsor_handling. Keeping sponsored polls is worse on the eve (driven by 2022, -0.032) and better four weeks out (+0.004); neither standpoint passes. Dropping them would move today's South Dakota race 27.5 points (its only polls are sponsored). Recommendation: no change; if this is revisited, test a sponsor-lean term rather than exclusion.
- pollster_house_effects, house_fundamentals_in_national, house_nat_swing. Effects at noise level (|gain| under 0.002; today's median margin shift 0.04 to 0.56). Recommendation: no action; they fail the majority rule by a coin's width.
Contradictions found
Opposite conclusions about the same quantity (the check in 2.6; seven quantities):
- Mean polling bias. Corrected three times in the published path (the scale intercept, -1.744;
the generic-ballot bias correction in N3, used by the House model and inside the statewide
ground-up prior; and the ground-up blend, whose gain the gate attributes to "correcting the 2020
polling miss") while
mean_bias_correction.mdandbias_centre.mdconcluded it must not be, and the methodology page tells readers "the lean is therefore reported, not subtracted" (src/elab/registry/calibration.py:CORRECTION_ATTEMPTS). The public sentence is inaccurate while the intercept is applied; correcting it is a communications fix the coordinator should make whatever the model decision. - Generic-ballot bias:
national_environment.md(supported) againsthouse_generic_ballot_bias.md(refused), the same day, neither citing the other. - An explicit fundamentals/poll blend: ADR-0007 rejects the knob; the sparse blend and the statewide ground-up blend are both such blends, the latter on top of a fit that already uses the fundamentals prior (prior information possibly counted twice).
- Design effect: 1.6 in code; measured near one (
poll_dispersion*.json). The audit agrees with the measurement (removing it does not hurt). - Sparse-poll error inflation: shipped as a fix (FM-104) after the same change was refused as challenger c9; the audit reproduces the refusal on log loss (0 of 3 cycles) while coverage favours the fix.
- Fitting at three polls: c6 passed; the published path still requires four.
- Publishing unpolled races from a prior: FM-140 publishes them; c18 refused an unpolled prior.
Inconsistent estimates that are not opposite conclusions (sigma_nat quoted as 2.60, 2.49, 2.48 and 2.44 in different files; sigma_race as 4.94 to 5.23; the realised sparse-poll error as 8.44, 8.92 and about 9.98) are recorded in the audit's working notes and are not flagged by the check.
Correctness findings under policy (a)
- House: the uncontested-seat rule gave LA-2, a Democratic seat, to the Republican because
neither candidate's quality record was known (
forecast_house_ground_up.py,fixed_d). That is a check of an adjustment's application, not a specification question. Follow-up (2026-10-05): the data side had already been repaired by FM-155 (LA-2 has been simulated, D >99%, since 2026-10-03), so this note was stale for the published number; the rule itself still defaulted a seat with no nominee on record to the Republican and is fixed by FM-173 (section 5). - Texas governor, re-derived step by step with
--check-race US-TX governor: reproduced -5.99 against the stored -5.97 (P 0.181 both). The scale intercept moves it 1.40 points; sponsored polls 0.36; the ground-up blend 0.41. The registered adjustments were applied correctly; the number's size is a specification question (item 1), not a defect.
Could not be tested, and why
poll_selection_asof (a knowability rule: switching it off is leakage); lv_rv_handling (nothing is applied; 64% of historical polls record no population); min_polls_rule (a threshold with no off state); prior_only_publication (switching it off unpublishes 12 races rather than changing them); chamber_correlation (one chamber outcome per cycle); house_uncontested_rule (the gate assigned those seats from actual winners); and the inactive or rejected entries (shock rules, district polls, the mean-bias correction itself).
5. Verified-bug fixes during the freeze
- FM-173 (2026-10-05, before the freeze): House one-party rule. A seat with no nominee on record
for either side was an automatic Republican seat; it is now simulated from its lean
(
split_one_party, shared by the published forecast, the audit and the counterfactual script). Effect today: none (no such seat). On the data that produced the LA-2 defect: 242.90 -> 243.89 Democratic seats, P(D majority) 0.980 -> 0.984.