What it is supposed to do
The specification, and a row per requirement saying where it is implemented and what demonstrates that it works.
REQUIREMENTS
0. What this system is
An election measurement and forecasting laboratory for U.S. federal and gubernatorial elections. It produces two distinct, never-conflated quantities per race, plus chamber-level distributions, and it carries the machinery to prove — on data it has never seen — that those quantities are better than simple alternatives.
It is not a clone of any existing forecaster. Existing forecasters are benchmarks and sources of testable hypotheses, never ground truth.
1. The two products
R1.1 NOWCAST — "if voting happened today"
For every in-scope race, a posterior distribution over the result conditional on current latent public opinion, with no future-drift uncertainty.
The electorate is specified (added after adversarial review, AR-01): the nowcast estimates the result if ballots were cast today by the electorate projected for Election Day, not by whoever would happen to turn out today. Likely-voter screens are pollsters' attempts to model the November electorate, so this is the quantity the poll data is evidence about. The alternative estimand — today's actual electorate — is a legitimate question this system does not answer, and the methodology page says so. See STATISTICAL_MODEL §11.0.
Must include: expected two-party and multi-candidate vote share; expected margin; credible intervals at 50/80/95%; win probability per candidate; and the implied chamber-composition distribution if every race were held today.
Retains uncertainty from: sampling error, systematic/correlated polling error, house-effect estimation error, likely-voter-screen error, turnout composition, undecided allocation, and model uncertainty.
Excludes: opinion movement between today and Election Day.
R1.2 ELECTION-DAY FORECAST — "what will happen in November"
The same objects, but marginalising over future opinion drift between today and Election Day, where the drift distribution is estimated from historical trajectory volatility at matched horizons, not assumed.
R1.3 Separation is structural, not cosmetic
The two must be produced by distinct code paths with distinct persisted artefacts and distinct API fields. No endpoint, table, chart, or report may present one where the other is implied. A test asserts that election-day predictive variance ≥ nowcast predictive variance for every race at every horizon > 0 days.
2. Scope
In scope (by phase, see ROADMAP.md)
- U.S. Senate (Phase 1 single race → Phase 5 all races)
- U.S. House, all 435 districts (Phase 6)
- Governors (Phase 6+)
- President — national, state-level, Electoral College (Phase 6+, cycle-dependent)
Explicitly out of scope for v1
State legislative races, ballot initiatives, primaries (modelled only as features for general elections, not forecast products), non-U.S. elections, turnout microtargeting, and any individual-level voter-file work.
Out of scope permanently
Proprietary source code, paywalled data used without authorisation, scraped content where a lawful stable alternative exists, and any design intended to favour a party or candidate.
3. Functional requirements
| ID | Requirement |
|---|---|
| F-01 | Ingest polls from authorised sources into an immutable raw store with full provenance (source, URL, retrieval ts, publication ts, raw bytes, parser version). |
| F-02 | Normalise raw polls into a typed schema capturing the field list in §5 of the build spec, with NULL distinguished from "not applicable" and from "not yet extracted". |
| F-03 | Detect and link duplicate, tracking and shared-sample polls, so that correlated observations are never treated as independent. |
| F-04 | Estimate pollster house effects and pollster-specific excess variance hierarchically, with uncertainty that is large when a pollster's history is thin. |
| F-05 | Fit a latent-opinion model per race producing a full posterior trajectory, not a point average. |
| F-06 | Fit a fundamentals model whose every retained feature has demonstrated incremental out-of-sample value. |
| F-07 | Simulate all races jointly under an empirically estimated correlated, heavy-tailed error structure. |
| F-08 | Produce and persist NOWCAST and ELECTION-DAY distributions per race and per chamber. |
| F-09 | Run strict walk-forward backtests where any forecast dated X uses only information knowable on X. |
| F-10 | Score forecasts with the metric suite in §19 of the build spec, sliced by cycle, office, geography, horizon, polling density, and competitiveness. |
| F-11 | Maintain baseline models and refuse to promote a complex model that cannot beat them out of sample. |
| F-12 | Operate a champion/challenger registry with pre-registered promotion criteria and an audit trail. |
| F-13 | Reproduce any archived forecast bit-for-bit (or within a declared numerical tolerance) from archived data + code + config + seeds. |
| F-14 | Attribute every forecast change to identifiable causes (new polls, ageing, house-effect revision, fundamentals change, model version change). |
| F-15 | Generate an automated post-election forensic error decomposition once certified results exist. |
| F-16 | Expose a machine-readable per-race forecast contract (§73 of build spec) via API. |
| F-17 | Provide a scenario engine whose outputs are labelled hypothetical and cannot write to production forecast tables. |
| F-18 | Surface a "data current through" timestamp and fail loudly on stale or failed ingestion. |
4. Non-functional requirements
| ID | Requirement |
|---|---|
| N-01 | Reproducibility. Every forecast records code commit, model version, dataset snapshot id, config hash, dependency lockfile hash, and RNG seeds. |
| N-02 | Auditability. A technically sophisticated outsider can trace any number to its inputs without reading the source code. |
| N-03 | Immutability of raw data. Raw records are append-only; corrections create revisions, never overwrites. |
| N-04 | Time-travel correctness. Every query used in backtesting is as_of-parameterised and bitemporal (valid time vs. transaction time). |
| N-05 | Neutrality. No component may be tuned toward a partisan outcome. A standing test measures signed forecast bias by party and flags drift. |
| N-06 | Honest uncertainty. Where evidence is thin, intervals widen. The system must be able to output "insufficient evidence". |
| N-07 | Performance envelope. On 16 threads: a single-race Bayesian fit <= 5 min; a full Senate cycle fit <= 60 min; 50,000 correlated chamber simulations <= 2 min; a 10-cycle Senate walk-forward backtest <= 12 h. House targets set in Phase 6 after profiling. |
| N-08 | Security. All fetched content is untrusted data. No secrets in git. See SECURITY.md. |
| N-09 | Portability. No dependence on machine-specific paths; the stack must stand up on a larger host from a lockfile and a config file. |
5. Explicit anti-requirements
The system must not:
- allow an LLM to emit a numeric probability, margin, or weight that reaches a forecast;
- silently reconcile conflicting sources;
- silently substitute synthetic data for real data in any non-test path;
- report a single point estimate without its uncertainty;
- present categorical ratings (Lean/Likely/Toss-up) as primary outputs;
- claim a component works before its tests pass;
- fit a decay curve, LV premium, correlation matrix, or tail parameter by assertion where it could be estimated.
6. Acceptance: Definition of Done per milestone
A milestone is done only when all hold:
- Implementation exists and is committed.
- Tests pass, including at least one synthetic-DGP recovery test where applicable.
- Provenance is populated for every record the milestone ingests.
- Walk-forward evaluation ran and its numbers are recorded in
experiments/. - Output is reproducible from archived inputs (N-01).
- Statistical assumptions are written down in
STATISTICAL_MODEL.md. - Model diagnostics (R-hat, ESS, divergences, posterior predictive checks) are within declared thresholds.
- An adversarial review has been performed and its findings recorded and answered.
- Known limitations are documented.
"The UI looks good" and "the API returns numbers" are not acceptance criteria.
7. Traceability, as of 2026-09-25
Where each requirement is implemented and what demonstrates it. Partial means the requirement is met in part and the remainder is named rather than rounded up.
| ID | Status | Where it lives | What shows it works |
|---|---|---|---|
| F-01 | met | elab.provenance.fetch, elab.provenance.archive, raw_artifact |
every poll and result traces to archived bytes with a sha256; tests/unit/test_fetch_guards.py |
| F-02 | met | elab.parse.*, poll, poll_question |
parser versions stored per row; golden tests in tests/golden/ |
| F-03 | met by a different mechanism, with one refusal and one unavailable source, both recorded | elab.models.fundamentals.environment, elab.models.pollster, fusion merging in elab.ingest.results, experiments/effective_n.md |
The requirement is that correlated observations are never treated as independent, and they are not — by two mechanisms rather than by the poll_link table originally envisaged. National environment: each firm is one cluster whose precision is capped at its largest single wave (independent_n), plus a firm-level jackknife. Measured need: 2020's generic-ballot window printed an effective n of 360,501 where Kish gives 695,887 and firm-clustering gives 26,032 — a factor of 27, and a margin sd of 0.62 points rather than 0.12. Race-level models: a per-pollster house effect (house_raw/house_sd) absorbs the correlated component of a firm's repeated polls within a race; a house effect is not identifiable on one national contest, which is why the generic ballot needed clustering and a state race does not. Duplicates and fusion lines are handled (FM-64, FM-70) and ADR-0014 covers cycle-level correlation. Refusal, measured: 538's tracking flag is deliberately not adopted — worth 1.0–2.6× on independent_n while pricing only panel overlap, when firm-capping prices house-effect concentration too and the jackknife shows that concentration is the larger problem (2022: analytic sd 0.59 against a firm jackknife of 2.25, Morning Consult holding 54.2% of the weight at D+3.19). Unavailable: poll.tracking_group_id, sample_group_id, is_tracking and poll_link.tracking_overlap hold zero rows, because no permitted source supplies a panel or tracking indicator for race-level polls — Wikipedia's race tables do not carry one, and the only enabled source that does (the 538 generic-ballot CSVs) feeds fundamental_series, which has no race_id and therefore cannot populate a poll column even in principle. The columns stay declared rather than dropped, so that a source which does supply it can be ingested without a migration; docs/DATABASE_SCHEMA.md now says they are empty instead of describing them in the present tense. |
| F-04 | met | elab.models.pollster |
experiments/house_effects.md; pooling was tried as a challenger and rejected (batch03_pooled_house.md) |
| F-05 | met | elab.models.latent.single_race |
experiments/m19_recovery_study.md (parameter recovery), diagnostic gates in passes_gates |
| F-06 | met | elab.models.fundamentals.prior |
every retained feature has a rejected alternative: house_prior_too_wide.md, house_prior_term_length.md, district_baseline.md |
| F-07 | met | elab.simulate.chamber, elab.models.correlation.structure |
correlated distribution is 3.5× wider than independent. The posterior predictive check on cross-race error correlation now passes on all six held-out cycles (experiments/phase5_validation_2026-09-29.json, host-tagged, two runs bit-identical): can the fitted structure produce the cycle that happened? 2014 reads tail p = 0.089 where it failed at 0.037, because sigma_nat fitted before 2014 is 2.05 rather than 1.21 once 1998–2002 are in the history, and the failing run predated that extension. 2014 is the marginal cycle against a next-closest 0.110. Seat coverage is 6 of 6 at a nominal 80%, which at n=6 is over-coverage rather than calibration, and the chamber PIT mean is 0.311 against 0.5 — the same Democratic lean the race scoring reports. sigma_region fits at 0.00, and that is now a bounded claim rather than a shrug: measured on the real skeleton, the estimator finds a true 2.0-point regional term 98% of the time and a 1.0-point one 38% of the time, so a regional term of 2.0 points or more is ruled out and one of 1.0 or less is undetectable here. The forecast's own regional signal, debiased for the noise on each regional mean, is 0.72 points across all races and 0.00 among well-polled ones — below the floor, so the fitted zero and the slice's 3.33-point spread agree. A spread over four groups is about twice a standard deviation and each regional mean carries ±1.2 of noise, which is where the apparent contradiction came from (experiments/phase5_exit_met.md, experiments/region_tension.md). |
| F-08 | met | elab.simulate.persist, race_forecast, chamber_forecast |
both products for Senate, House, governors and the electoral college |
| F-09 | met | elab.repo.* (as_of on every read) |
tests/leakage/, oracle canaries, docs/HOLDOUT_LEDGER.md |
| F-10 | met, with one slice empty for want of data | scripts/score_stored_forecasts.py, elab.evaluate.scored |
All six slices the requirement names — cycle, office, geography, horizon, polling density, competitiveness — are built on one truth join and persisted to model_calibration as …#slice=value rows. This row previously claimed four of six and admitted two; both halves were wrong, because the scorer sliced by cycle only and the others lived in forensics.py and diagnose_calibration.py, the first of which used its own SQL without MIN_TWO_WAY_SHARE or the third-candidate-won refusal. Findings in experiments/scoring_slices.md: the 20+-point bucket is the worst forecast (bias +3.65, 90% coverage 0.848), densely polled races are over-covered (0.954 against a nominal 0.80), and a 3.3-point bias spread across census regions sits on the partition where sigma_region estimates to 0.00. The horizon slice is built and reports nothing: every scored forecast stands at election day, because no multi-horizon backtest is persisted as scored forecasts. That is a data gap, not a missing slice, and it is stated rather than rounded up. |
| F-11 | met | elab.models.baselines, elab.registry.promotion, scripts/diagnose_calibration.py |
The champion beats the baseline out of sample: MAE 4.89 against a recency- and size-weighted polling average's 5.51, over 435 election-day forecasts, one per race, across ten cycles 2006-2024, and it reduces the polls' lean rather than adding to it (+2.08 against +2.86). Better in seven of the ten cycles; worse in 2006 (n=4), 2008 and 2014. Realised coverage 0.471 / 0.798 / 0.903 against a nominal 0.50 / 0.80 / 0.90 (single_race_latent/0.7.0). An earlier draft of this row cited 5.39 against 6.04 over 900 points; that set was 0.6.0 mid-republication and counted every republished race twice, once under each specification, so it averaged pre-fix and post-fix forecasts of the same race. One forecast per race is the comparison. Thirteen challengers evaluated, two promoted, eleven refused with their numbers recorded (experiments/README.md). This row previously cited only the challenger count and no baseline comparison at all, which is how an unsupported claim that the model lost to the polling average stood for part of a day and set a work programme: diagnose_calibration.py defaulted to --semver 0.3.0, three versions stale, and nobody had asked the champion (FM-105). |
| F-12 | met | model_version, promotion_audit, experiment |
pre-registrations in experiments/preregistration_*.md; refusals recorded with promotions |
| F-13 | met | scripts/reproduce_forecast.py, elab.simulate.seeds |
reproduces a Senate race forecast from a clean restore within stated Monte Carlo tolerance. This row was false for the House until 2026-09-27: forecast_house.py seeded each district from abs(hash(geo)) % 2**32, and Python salts string hashing per process, so the seat distribution could not be reproduced from its own recorded inputs and the seed was not recorded because it was not knowable (FM-95). reproduce_forecast.py passed throughout, because it reproduces a Senate forecast, which takes its seed from an argument. Seeds now come from stable_seed, a blake2b digest identical across processes; two consecutive House runs give the same seat distribution. This row was false from 2026-09-26 until 2026-09-29 and said met throughout. reproduce_forecast.py passed lv_sd, undecided_sd and model_sd, the three placeholders 0.4.0 deleted when sigma_race replaced them, so it raised KeyError on every run published since and knew nothing of the prior, the scale correction or the measured sparse-poll sd added after. Nobody ran it; I re-asserted the row today by checking the file existed (FM-110). Fixed and verified: a well-polled 2026 Senate forecast and a sparse one both rebuild with margin delta 0.000000 and win-probability delta 0.000000, against tolerances of 0.129 and 0.252 points. tests/integration/test_reproduce_forecast_is_current.py pins the argument list against the signature so it cannot rot for a fourth version. |
| F-14 | met | forecast_attribution (migration 0023), scripts/attribute_forecasts.py |
answered on the race page; needs known_at, which is why it was once refused. Was silently broken for three specification versions and found the day the scheduler was installed: the replay passed variance terms 0.4.0 deleted and omitted the prior and scale correction added since, so it would have reported a specification change as movement in the race (FM-92). |
| F-15 | met | scripts/forensics.py |
one file per cycle in experiments/forensics_*.md; found the defect behind c9 |
| F-16 | met | GET /races/{id}/forecast |
full contract: quantiles, win probabilities by candidacy, variance decomposition, provenance |
| F-17 | met | scripts/publish_chamber.py --shift, forecast_run.is_scenario |
tests/unit/test_scenario_isolation.py parses the API's SQL and asserts every read excludes scenarios |
| F-18 | met | /health, /model-health, elab.obs.health |
source anomalies measured against each source's own history; the publication gate refuses stale data |
| N-01 | met | forecast_run.inputs, elab.simulate.seeds |
git commit, lockfile hash, seeds and config hash per run. The seeds clause was unsatisfiable for the House until 2026-09-27 — the per-district seed came from hash(), which is salted per process, so it did not exist until the run began (FM-95). A requirement that a code path cannot satisfy is evidence about the code path, and this row said met without anyone asking which forecast had been reproduced. This row was false from 2026-09-26 until 2026-09-29 and said met throughout. reproduce_forecast.py passed lv_sd, undecided_sd and model_sd, the three placeholders 0.4.0 deleted when sigma_race replaced them, so it raised KeyError on every run published since and knew nothing of the prior, the scale correction or the measured sparse-poll sd added after. Nobody ran it; I re-asserted the row today by checking the file existed (FM-110). Fixed and verified: a well-polled 2026 Senate forecast and a sparse one both rebuild with margin delta 0.000000 and win-probability delta 0.000000, against tolerances of 0.129 and 0.252 points. tests/integration/test_reproduce_forecast_is_current.py pins the argument list against the signature so it cannot rot for a fourth version. |
| N-02 | met | docs/, experiments/README.md |
thirty-nine experiment files indexed; every screen number carries a run_id |
| N-03 | met | append-only raw_artifact, supersession everywhere |
corrections create revisions (FM-48, FM-55) |
| N-04 | met | bitemporal recorded_at/superseded_at |
tests/leakage/; the one relaxation is stated in FM-56 |
| N-05 | met | tests/unit/test_neutrality.py |
signed bias measured from the stored record and printed. +2.40 points over the 295 races scored under the current champion 0.6.0, down from +3.60 under 0.5.0: the promoted scale correction removed a third of it (c15, 2026-09-27). What remains sits in the cycle means — +5.03 in 2016 and +4.99 in 2020 against −1.27 in 2018 — which is the national miss sigma_nat exists to carry, and nothing estimable before a cycle predicts it. |
| N-06 | met | refusals throughout | withheld panels, "insufficient evidence" paths, UnjustifiedUncertainty |
| N-07 | met, and all four clauses now measured | scripts/time_walk_forward.py, scripts/time_simulation.py, experiments/performance_envelope.md |
On Skynet10, the 16-thread host the requirement specifies: single-race fit 4–14 s (target 5 min); full Senate cycle 311 s (target 60 min), timed end to end for the first time on 2026-09-27; 50,000 correlated chamber simulations in 0.07 s for the Senate and 4.59 s for 435 House districts (target 2 min), also timed for the first time; eleven-cycle Senate walk-forward 6.75 h serial / 1.79 h at 4 workers against a 12 h target — a projection from 38 timed fits over an exact 3,824-fit grid, not an end-to-end run, and the artifact says so. The House clause is now a target rather than a promise: 435-district simulation ≤ 2 min, full district forecast ≤ 10 min, set from measurement. This row previously read "met, measured" while two of the four clauses had never been timed and the timing artifact carried no host, which TEST_PLAN.md requires (FM-96). Practical figure: about four minutes on Skynet-Three at 64 workers (FM-83). |
| N-08 | met | docs/SECURITY.md, redact_url, header auth |
no secrets in git; prohibited sources disabled in source_registry |
| N-09 | met | scripts/bootstrap_host.sh, elab.config.settings, deploy/render-units.sh, tests/unit/test_portability.py, experiments/second_host.md |
Done, on 2026-09-27: the stack stands up on Skynet-Three (dual EPYC 7742) from the lockfile and a config file — migration head 0029, 35 tables, the four ARCHITECTURE §6 roles, 25 sources seeded, and 1,014 of 1,065 tests passing with every skip self-explaining (the laptop collects the same 1,065 and passes 1,064). bootstrap_host.sh is idempotent and needs no root: a user-space PostgreSQL from conda-forge via micromamba, socket-only, per ADR-0002. Attempting it is what found the reason this row was partial for months (FM-97): PyMC was in no lockfile, so a fresh host got a stack that could not fit a model and every synthetic-DGP recovery test — Definition of Done criterion 2 — failed; there was no requirements-dev.txt, so the test environment was not reproducible either; .env.example carried three real home paths and the portability test could not see it, because .example was not in its suffix list; and pguser defaulted to one developer's username. All four fixed. Not established: that both hosts produce identical forecasts from the same inputs — the nearest evidence is FM-91's cross-path comparison, mean difference 0.039 points against 0.055 of seed noise. |