FutureBallotU.S. election forecasts, with every number traced to its source

How it is tested

The walk-forward protocol, the leakage guards, and what counts as evidence that a change is an improvement rather than a coincidence.

BACKTESTING

The backtest harness is the most important component in the repository. A forecasting model without a leak-proof evaluation harness is an unfalsifiable claim, and unfalsifiable claims are what this project exists to avoid.

1. The contract

A forecast dated X may use only information that was knowable on X.

"Knowable" is stricter than "about a period before X". A poll fielded in March but published in April is not knowable in March. A GDP figure for Q1 2016 as revised in 2019 is not knowable in 2016. A pollster rating computed after the 2020 election is not knowable in October 2020.

2. How the contract is enforced

2.1 Structurally, not by discipline

There is exactly one way to read modelling data: the repository layer, whose every method requires an as_of parameter (ARCHITECTURE §3). The filter is always:

WHERE recorded_at <= :as_of
  AND (superseded_at IS NULL OR superseded_at > :as_of)

on transaction time, plus a source-specific knowability filter on valid time, e.g. for polls published_at <= :as_of, for FEC filings filing_date <= :as_of, for economic series vintage_date <= :as_of.

Production is as_of = now(). There is no separate production query path that could drift from the backtest path.

2.2 Tests that enforce it

Test Assertion
test_no_raw_sql_outside_repo No module outside src/elab/repo/ executes SQL against L1–L5 tables
test_repo_methods_require_as_of Every public repository method signature has a non-default as_of
test_bitemporal_columns_present Every L1–L5 table has recorded_at and superseded_at

Revised after adversarial review (AR-09). An earlier draft asserted that the visible row set at T1 is a subset of that at T2. That is false by design: a row superseded between T1 and T2 is visible at T1 and invisible at T2. The original test would have failed against a correct implementation and would then most likely have been "fixed" by weakening the bitemporal logic — which is how a real guarantee gets quietly deleted. The two assertions above are both true and both catch the leak the original was aiming at. | test_asof_no_future_records | No row visible at T1 has recorded_at > T1 | | test_asof_recording_monotone | For T1 < T2, rows ever recorded by T1 are a subset of those ever recorded by T2 |

2.3 The canary test

The strongest leakage detector is an oracle test. Insert a synthetic poll dated after the as-of date whose result is wildly different from the truth. Run the backtest. If the forecast moves at all, information is flowing backwards. This runs in CI on every commit that touches repo/, backtest/, or any model.

A second canary: shuffle the outcomes of held-out races while leaving all features intact. A model with no leakage should score at chance on the shuffled set. Scoring better than chance means an outcome-derived quantity reached the features.

2.4 Leakage taxonomy, with the detector for each

# Leak Detector
L1 Future polls published_at <= as_of filter + canary test
L2 Revised economic data vintage_date in the PK; test that no series has a single vintage
L3 Post-election pollster ratings Pollster-quality model is refitted per as-of date from data ≤ as-of; test that its parameters differ across as-of dates
L4 Post-hoc district classifications Boundary and rating features keyed by boundary_id and effective date
L5 Finalised candidate fields candidacy rows are bitemporal; a candidate who withdrew in October is present in June
L6 Corrected datasets Corrections supersede, never overwrite; the old value remains visible to earlier as-of dates
L7 Final turnout data Turnout series carry vintage dates; election-year turnout is knowable only after the election
L8 Post-election demographic estimates ACS vintages; exit-poll-derived features are banned as inputs to pre-election forecasts
L9 Survivorship in the pollster set Pollsters that quit are present in the as-of-dated pollster table; test that defunct pollsters appear in backtests dated before defunct_at
L10 Hyperparameters tuned on the test cycle The holdout ledger (§6); pre-registration of experiments
L11 Author leakage — the developer knows 2016 was a polling miss and designs the error structure accordingly Partially irreducible (FM-30). Mitigated by pre-registration and by final evaluation on cycles the design predates
L12 Entity-resolution using future information (e.g. knowing two pollster brands merged) Alias tables are bitemporal

L11 deserves emphasis: it cannot be fully eliminated. Anyone building this model in 2026 knows what happened in 2016, 2020, and 2022. Estimating a fat-tailed correlated error structure "from the data" when you already know the answer is a softer form of leakage that no code can detect. The only honest mitigations are (a) pre-registering the structure before looking at per-cycle results, (b) reporting how much of the tail behaviour is prior versus likelihood, and (c) treating future real elections as the only truly clean test.

3. Walk-forward protocol

for cycle in cycles:
    for as_of in forecast_dates(cycle):        # e.g. every 7 days from -540 to -1
        snapshot = repo.snapshot(as_of)
        model    = fit(snapshot, config)       # refit, not reused across as_of
        nowcast  = simulate_nowcast(model)
        forecast = simulate_election_day(model)
        persist(run_id, as_of, cycle, nowcast, forecast)
score(all_runs, certified_results)

Two disciplines that are easy to get wrong:

3.1 Compute budget (added after adversarial review, AR-05)

The naive protocol does not fit on this machine, and the arithmetic says so plainly: ~78 weekly as-of dates × 10 cycles ≈ 780 full fits per specification. At the 60-minute target for a joint Senate fit that is 780 hours — 32 days — for one specification. With several specifications and an automated challenger generator, the design as first written cannot run.

Four measures bring it into range:

  1. Tiered as-of grids. Weekly inside the final 120 days (where forecasts change and where most evaluation interest lies), fortnightly out to 365 days, monthly before that. Fits per cycle fall from ~78 to ~30. The grid is configuration, and the tiering is reported with every result so nobody mistakes a coarse early grid for a dense one.

  2. Warm starts along the time axis. Initialising the fit at as-of T from the posterior at the previous as-of date is not leakage — that date precedes T — and it cuts warmup substantially. It must be validated: a stratified subsample is fitted both warm and cold and the posteriors compared, with the discrepancy recorded as an error budget. If warm starts bias results, they are dropped.

  3. A fast inner loop with a measured approximation cost. Laplace or ADVI approximations across the bulk of the grid, with full NUTS at a stratified subsample of as-of dates. The approximation error is quantified against those NUTS anchors and reported alongside every backtest number. This is a documented cost, never a silent one — an approximate backtest that presents itself as exact is a way of lying with arithmetic.

  4. Screening before full evaluation. A challenger earns a full walk-forward only after surviving one or two cycles. This conserves compute and holdout (ADR-0010), which are both scarce and which deplete together.

Revised estimate: ~30 as-of dates × 10 cycles, mostly approximate, ≈ 1–2 days per full specification evaluation. That is the difference between a project that can iterate and one that cannot.

This is also why Phase 2 builds the harness before the expensive model exists, and why ARCHITECTURE §5 caps worker memory: a backtest that swaps this machine takes a week instead of a day.

3.2 The effective sample size of the evaluation (added after AR-02)

Refitting everything at each as-of date has a consequence the first draft left unstated. The systematic-error scale σ_nat can only be estimated from cycles before the as-of date. A forecast dated 2004 has perhaps three prior cycles to learn it from, so its σ_nat posterior is close to pure prior, and that backtest year tests the prior rather than the model.

Therefore:

Anyone reporting "calibrated across ten cycles" from this design would be overstating the evidence by roughly a factor of two.

4. Validation designs

Design Purpose
Rolling-origin Primary. Train on cycles < c, evaluate on cycle c, for each c
Leave-one-cycle-out Robustness across eras; detects a model that depends on one cycle
Geographic holdout Hold out whole regions; detects geography-specific overfit
Office holdout Fit on Senate, evaluate on Governor; tests structural transfer
Untouched terminal holdout The most recent complete cycle, sealed (§6)

5. Metrics

Probabilistic: Brier score, log score, continuous ranked probability score (CRPS) on the margin, calibration curves with bootstrap bands, expected calibration error, and credible-interval coverage at 50/80/95%.

Point: margin MAE and RMSE, signed bias (overall and by party, per N-05).

Aggregate: seat-count MAE, seat-distribution calibration (is the realised seat count inside the 80% interval 80% of the time, across cycles?), chamber-control probability calibration.

Tail: performance restricted to races where the realised outcome fell outside the 90% interval; frequency of such events versus the nominal 10%.

Every metric is reported with a standard error. With ~35 Senate races per cycle and ~10 cycles, the effective sample for chamber-level claims is ten, and cross-race correlation means the effective sample for a cycle-level claim is closer to one. A difference in Brier score between two models is usually noise, and the harness says so rather than ranking silently.

5.1 Slicing

All metrics are computed sliced by: cycle, office, region, forecast horizon bucket, polling density (0 polls / 1–3 / 4–10 / 10+), competitiveness (|margin| bands), and pollster-mix. A model that is well calibrated overall and badly calibrated in sparse-polling districts is a model that will get the House wrong. The slice report is a gate, not an appendix.

6. The holdout ledger

Holdout is a consumable resource. Every look at a held-out cycle spends some of it.

experiments.holdout_consumed records, per experiment, which cycles were evaluated and how many times. docs/HOLDOUT_LEDGER.md is generated from that table and states, for each cycle:

Policy:

This is the mechanism that stops the champion/challenger loop from degenerating into an elaborate overfit of ten historical elections — which, absent this ledger, is exactly what a tireless automated researcher would produce.

7. What counts as an improvement

A challenger beats the champion only if:

  1. Out-of-sample log score improves by more than the paired bootstrap standard error of the difference, computed with cycle-level resampling (races within a cycle are not independent); 1b. Multiplicity is controlled across the challenger batch (added after adversarial review, AR-11). Challengers are pre-registered in batches before evaluation; Benjamini–Hochberg FDR control is applied across the batch at a target FDR set in configuration. Per-comparison error control, which the first draft specified, guarantees a steady trickle of false promotions once a generator is producing challengers continuously — the null-challenger test (FM-32) would have detected the symptom while nothing in the design prevented the cause;
  2. Calibration does not deteriorate on any slice in §5.1;
  3. The improvement holds in at least two held-out cycles, or one cycle plus one held-out geography;
  4. Tail coverage does not deteriorate;
  5. There is a stated mechanism for why it improves — an unexplained improvement is treated as a suspected leak until proven otherwise;
  6. Complexity cost is acceptable and documented;
  7. An adversarial review has attempted to break it and failed.

Failure to meet any of these means the champion stays. "The challenger looked better on average" is not a promotion criterion.

7.1 Implementation (elab.registry.promotion)

evaluate_batch takes a pre-registered batch of challengers with per-cycle paired score differences and returns a decision each. Batch composition must be fixed before results are seen: dropping the losers and correcting over the survivors controls nothing.

The p-value comes from resampling cycles, with a +1 correction so that a p-value of exactly zero is unattainable — with seven cycles it is not evidence that strong, and letting the arithmetic say otherwise invites the ratchet. Benjamini–Hochberg then bounds the proportion of promotions that are false, which is the quantity that matters when the aim is a sequence of good decisions rather than a single one.

A challenger is promoted only if all of: q ≤ target FDR; the mean difference actually favours it; it replicated on at least two folds; it states a mechanism; and its complexity cost is within budget. The mechanism requirement is not bureaucracy — leaks look exactly like improvements, so an unexplained gain is treated as a suspected leak until shown otherwise. In testing, a challenger with q = 0.001 was rejected for having no stated mechanism, which is the intended behaviour.

promote() is the only path to changing a champion. It refuses a decision the gate rejected and refuses to run without an adversarial review id, re-checking at the last point before the champion changes: a promotion that bypassed the gate would be indistinguishable afterwards from one that passed it. Retirement and promotion happen in one transaction because the schema permits only one champion per family.

The gate is calibrated, not merely specified. test_null_challengers_are_promoted_at_about_the_nominal_rate feeds it 1,200 challengers that are the champion plus noise and asserts the promotion rate does not exceed the nominal FDR. A gate that cannot be shown to resist promoting noise is worse than no gate, because it launders noise as evidence (FM-32).