How it is tested
The walk-forward protocol, the leakage guards, and what counts as evidence that a change is an improvement rather than a coincidence.
BACKTESTING
The backtest harness is the most important component in the repository. A forecasting model without a leak-proof evaluation harness is an unfalsifiable claim, and unfalsifiable claims are what this project exists to avoid.
1. The contract
A forecast dated
Xmay use only information that was knowable onX.
"Knowable" is stricter than "about a period before X". A poll fielded in March but published in April is not knowable in March. A GDP figure for Q1 2016 as revised in 2019 is not knowable in 2016. A pollster rating computed after the 2020 election is not knowable in October 2020.
2. How the contract is enforced
2.1 Structurally, not by discipline
There is exactly one way to read modelling data: the repository layer, whose every method
requires an as_of parameter (ARCHITECTURE §3). The filter is always:
WHERE recorded_at <= :as_of
AND (superseded_at IS NULL OR superseded_at > :as_of)
on transaction time, plus a source-specific knowability filter on valid time, e.g. for polls
published_at <= :as_of, for FEC filings filing_date <= :as_of, for economic series
vintage_date <= :as_of.
Production is as_of = now(). There is no separate production query path that could drift from
the backtest path.
2.2 Tests that enforce it
| Test | Assertion |
|---|---|
test_no_raw_sql_outside_repo |
No module outside src/elab/repo/ executes SQL against L1–L5 tables |
test_repo_methods_require_as_of |
Every public repository method signature has a non-default as_of |
test_bitemporal_columns_present |
Every L1–L5 table has recorded_at and superseded_at |
Revised after adversarial review (AR-09). An earlier draft asserted that the visible row set at
T1is a subset of that atT2. That is false by design: a row superseded betweenT1andT2is visible atT1and invisible atT2. The original test would have failed against a correct implementation and would then most likely have been "fixed" by weakening the bitemporal logic — which is how a real guarantee gets quietly deleted. The two assertions above are both true and both catch the leak the original was aiming at. |test_asof_no_future_records| No row visible atT1hasrecorded_at > T1| |test_asof_recording_monotone| ForT1 < T2, rows ever recorded byT1are a subset of those ever recorded byT2|
2.3 The canary test
The strongest leakage detector is an oracle test. Insert a synthetic poll dated after the
as-of date whose result is wildly different from the truth. Run the backtest. If the forecast
moves at all, information is flowing backwards. This runs in CI on every commit that touches
repo/, backtest/, or any model.
A second canary: shuffle the outcomes of held-out races while leaving all features intact. A model with no leakage should score at chance on the shuffled set. Scoring better than chance means an outcome-derived quantity reached the features.
2.4 Leakage taxonomy, with the detector for each
| # | Leak | Detector |
|---|---|---|
| L1 | Future polls | published_at <= as_of filter + canary test |
| L2 | Revised economic data | vintage_date in the PK; test that no series has a single vintage |
| L3 | Post-election pollster ratings | Pollster-quality model is refitted per as-of date from data ≤ as-of; test that its parameters differ across as-of dates |
| L4 | Post-hoc district classifications | Boundary and rating features keyed by boundary_id and effective date |
| L5 | Finalised candidate fields | candidacy rows are bitemporal; a candidate who withdrew in October is present in June |
| L6 | Corrected datasets | Corrections supersede, never overwrite; the old value remains visible to earlier as-of dates |
| L7 | Final turnout data | Turnout series carry vintage dates; election-year turnout is knowable only after the election |
| L8 | Post-election demographic estimates | ACS vintages; exit-poll-derived features are banned as inputs to pre-election forecasts |
| L9 | Survivorship in the pollster set | Pollsters that quit are present in the as-of-dated pollster table; test that defunct pollsters appear in backtests dated before defunct_at |
| L10 | Hyperparameters tuned on the test cycle | The holdout ledger (§6); pre-registration of experiments |
| L11 | Author leakage — the developer knows 2016 was a polling miss and designs the error structure accordingly | Partially irreducible (FM-30). Mitigated by pre-registration and by final evaluation on cycles the design predates |
| L12 | Entity-resolution using future information (e.g. knowing two pollster brands merged) | Alias tables are bitemporal |
L11 deserves emphasis: it cannot be fully eliminated. Anyone building this model in 2026 knows what happened in 2016, 2020, and 2022. Estimating a fat-tailed correlated error structure "from the data" when you already know the answer is a softer form of leakage that no code can detect. The only honest mitigations are (a) pre-registering the structure before looking at per-cycle results, (b) reporting how much of the tail behaviour is prior versus likelihood, and (c) treating future real elections as the only truly clean test.
3. Walk-forward protocol
for cycle in cycles:
for as_of in forecast_dates(cycle): # e.g. every 7 days from -540 to -1
snapshot = repo.snapshot(as_of)
model = fit(snapshot, config) # refit, not reused across as_of
nowcast = simulate_nowcast(model)
forecast = simulate_election_day(model)
persist(run_id, as_of, cycle, nowcast, forecast)
score(all_runs, certified_results)
Two disciplines that are easy to get wrong:
- Everything is refitted at each as-of date, including pollster quality, house effects, fundamentals coefficients, decay parameters, and correlation structure. Reusing a once-fitted object across as-of dates is leakage L3/L10 in disguise.
- Hyperparameter selection happens inside the walk-forward loop, on data prior to the as-of date. Selecting a decay constant on all ten cycles and then "backtesting" it is not a backtest.
3.1 Compute budget (added after adversarial review, AR-05)
The naive protocol does not fit on this machine, and the arithmetic says so plainly: ~78 weekly as-of dates × 10 cycles ≈ 780 full fits per specification. At the 60-minute target for a joint Senate fit that is 780 hours — 32 days — for one specification. With several specifications and an automated challenger generator, the design as first written cannot run.
Four measures bring it into range:
-
Tiered as-of grids. Weekly inside the final 120 days (where forecasts change and where most evaluation interest lies), fortnightly out to 365 days, monthly before that. Fits per cycle fall from ~78 to ~30. The grid is configuration, and the tiering is reported with every result so nobody mistakes a coarse early grid for a dense one.
-
Warm starts along the time axis. Initialising the fit at as-of
Tfrom the posterior at the previous as-of date is not leakage — that date precedesT— and it cuts warmup substantially. It must be validated: a stratified subsample is fitted both warm and cold and the posteriors compared, with the discrepancy recorded as an error budget. If warm starts bias results, they are dropped. -
A fast inner loop with a measured approximation cost. Laplace or ADVI approximations across the bulk of the grid, with full NUTS at a stratified subsample of as-of dates. The approximation error is quantified against those NUTS anchors and reported alongside every backtest number. This is a documented cost, never a silent one — an approximate backtest that presents itself as exact is a way of lying with arithmetic.
-
Screening before full evaluation. A challenger earns a full walk-forward only after surviving one or two cycles. This conserves compute and holdout (ADR-0010), which are both scarce and which deplete together.
Revised estimate: ~30 as-of dates × 10 cycles, mostly approximate, ≈ 1–2 days per full specification evaluation. That is the difference between a project that can iterate and one that cannot.
This is also why Phase 2 builds the harness before the expensive model exists, and why ARCHITECTURE §5 caps worker memory: a backtest that swaps this machine takes a week instead of a day.
3.2 The effective sample size of the evaluation (added after AR-02)
Refitting everything at each as-of date has a consequence the first draft left unstated. The
systematic-error scale σ_nat can only be estimated from cycles before the as-of date. A
forecast dated 2004 has perhaps three prior cycles to learn it from, so its σ_nat posterior is
close to pure prior, and that backtest year tests the prior rather than the model.
Therefore:
- The effective evaluation set is not ten cycles. For calibration claims that depend on
σ_natit is roughly the most recent five or six; for correlated-error and chamber-control behaviour, fewer still. - Every backtest report states, per cycle, the share of
σ_nat's posterior precision coming from the likelihood versus the prior. - Headline calibration claims are restricted to cycles exceeding a stated likelihood-contribution threshold. Earlier cycles are still run and still reported — they are informative about the data pipeline and about point accuracy — but they are not cited as evidence that the uncertainty is calibrated.
Anyone reporting "calibrated across ten cycles" from this design would be overstating the evidence by roughly a factor of two.
4. Validation designs
| Design | Purpose |
|---|---|
| Rolling-origin | Primary. Train on cycles < c, evaluate on cycle c, for each c |
| Leave-one-cycle-out | Robustness across eras; detects a model that depends on one cycle |
| Geographic holdout | Hold out whole regions; detects geography-specific overfit |
| Office holdout | Fit on Senate, evaluate on Governor; tests structural transfer |
| Untouched terminal holdout | The most recent complete cycle, sealed (§6) |
5. Metrics
Probabilistic: Brier score, log score, continuous ranked probability score (CRPS) on the margin, calibration curves with bootstrap bands, expected calibration error, and credible-interval coverage at 50/80/95%.
Point: margin MAE and RMSE, signed bias (overall and by party, per N-05).
Aggregate: seat-count MAE, seat-distribution calibration (is the realised seat count inside the 80% interval 80% of the time, across cycles?), chamber-control probability calibration.
Tail: performance restricted to races where the realised outcome fell outside the 90% interval; frequency of such events versus the nominal 10%.
Every metric is reported with a standard error. With ~35 Senate races per cycle and ~10 cycles, the effective sample for chamber-level claims is ten, and cross-race correlation means the effective sample for a cycle-level claim is closer to one. A difference in Brier score between two models is usually noise, and the harness says so rather than ranking silently.
5.1 Slicing
All metrics are computed sliced by: cycle, office, region, forecast horizon bucket, polling density (0 polls / 1–3 / 4–10 / 10+), competitiveness (|margin| bands), and pollster-mix. A model that is well calibrated overall and badly calibrated in sparse-polling districts is a model that will get the House wrong. The slice report is a gate, not an appendix.
6. The holdout ledger
Holdout is a consumable resource. Every look at a held-out cycle spends some of it.
experiments.holdout_consumed records, per experiment, which cycles were evaluated and how
many times. docs/HOLDOUT_LEDGER.md is generated from that table and states, for each cycle:
- number of distinct model specifications evaluated against it;
- date of first evaluation;
- current status:
sealed/validation/burned.
Policy:
- The most recent complete cycle starts sealed. Unsealing requires an explicit, logged decision and it never re-seals.
- A cycle evaluated by more than 20 specifications is marked
burnedand may no longer be cited as evidence of out-of-sample performance. - When all cycles are burned, honest evaluation requires waiting for a real election. The system is allowed to say that.
This is the mechanism that stops the champion/challenger loop from degenerating into an elaborate overfit of ten historical elections — which, absent this ledger, is exactly what a tireless automated researcher would produce.
7. What counts as an improvement
A challenger beats the champion only if:
- Out-of-sample log score improves by more than the paired bootstrap standard error of the difference, computed with cycle-level resampling (races within a cycle are not independent); 1b. Multiplicity is controlled across the challenger batch (added after adversarial review, AR-11). Challengers are pre-registered in batches before evaluation; Benjamini–Hochberg FDR control is applied across the batch at a target FDR set in configuration. Per-comparison error control, which the first draft specified, guarantees a steady trickle of false promotions once a generator is producing challengers continuously — the null-challenger test (FM-32) would have detected the symptom while nothing in the design prevented the cause;
- Calibration does not deteriorate on any slice in §5.1;
- The improvement holds in at least two held-out cycles, or one cycle plus one held-out geography;
- Tail coverage does not deteriorate;
- There is a stated mechanism for why it improves — an unexplained improvement is treated as a suspected leak until proven otherwise;
- Complexity cost is acceptable and documented;
- An adversarial review has attempted to break it and failed.
Failure to meet any of these means the champion stays. "The challenger looked better on average" is not a promotion criterion.
7.1 Implementation (elab.registry.promotion)
evaluate_batch takes a pre-registered batch of challengers with per-cycle paired score
differences and returns a decision each. Batch composition must be fixed before results are
seen: dropping the losers and correcting over the survivors controls nothing.
The p-value comes from resampling cycles, with a +1 correction so that a p-value of
exactly zero is unattainable — with seven cycles it is not evidence that strong, and letting
the arithmetic say otherwise invites the ratchet. Benjamini–Hochberg then bounds the
proportion of promotions that are false, which is the quantity that matters when the aim is
a sequence of good decisions rather than a single one.
A challenger is promoted only if all of: q ≤ target FDR; the mean difference actually
favours it; it replicated on at least two folds; it states a mechanism; and its complexity
cost is within budget. The mechanism requirement is not bureaucracy — leaks look exactly
like improvements, so an unexplained gain is treated as a suspected leak until shown
otherwise. In testing, a challenger with q = 0.001 was rejected for having no stated
mechanism, which is the intended behaviour.
promote() is the only path to changing a champion. It refuses a decision the gate rejected
and refuses to run without an adversarial review id, re-checking at the last point before the
champion changes: a promotion that bypassed the gate would be indistinguishable afterwards
from one that passed it. Retirement and promotion happen in one transaction because the
schema permits only one champion per family.
The gate is calibrated, not merely specified. test_null_challengers_are_promoted_at_about_the_nominal_rate
feeds it 1,200 challengers that are the champion plus noise and asserts the promotion rate
does not exceed the nominal FDR. A gate that cannot be shown to resist promoting noise is
worse than no gate, because it launders noise as evidence (FM-32).