Decisions and why
The architecture decision records. Each one states the alternatives that were rejected and what would have to be true to revisit it.
DECISIONS (ADR log)
Each entry: the decision, the alternatives actually considered, why this one, and what would overturn it. An ADR with no stated falsifier is not a decision, it is a preference.
ADR-0001 · Two products, structurally separated
Decision. NOWCAST and ELECTION-DAY FORECAST are produced by distinct code paths and stored
in rows distinguished by a kind CHECK constraint.
Alternatives. One pipeline with a horizon parameter; a single "forecast" with a toggle.
Why. A parameterised single path makes conflation a one-line bug and makes it invisible in storage. The two answer different questions and their variance budgets differ by an entire term (STATISTICAL_MODEL §11). Separation at the schema level makes the mistake unrepresentable.
Falsifier. If the drift term turns out to be estimable as a pure post-hoc variance inflation with no structural difference, the paths could merge. Evidence needed: identical posterior predictive behaviour across all horizons.
ADR-0002 · PostgreSQL, installed into user space via micromamba
Decision. PostgreSQL 16 remains the system of record, installed from conda-forge via
micromamba into $HOME, running as the user on a local socket.
Alternatives. (a) sudo apt install postgresql — blocked: no passwordless sudo.
(b) Docker Postgres — blocked: no container runtime and none installable without root.
(c) SQLite as system of record — rejected: no role separation (ARCHITECTURE §6 depends on it),
weak constraint vocabulary, no EXCLUDE USING gist for boundary ranges, single-writer.
(d) DuckDB as system of record — rejected for the same reasons plus concurrency; adopted
instead as a read-side analytical accelerator over Parquet exports.
Why. The schema's integrity guarantees — exclusion constraints on boundary ranges, partial unique indexes for bitemporality, revoked grants for raw immutability, four-role privilege separation — are the mechanism by which several critical failure modes become impossible rather than merely detectable. Those guarantees are Postgres features, not preferences. Micromamba obtains them without root.
Falsifier. If micromamba's Postgres proves unstable or unmaintainable here, the fallback is to ask the user for a single privileged install command (a legitimate interruption per build spec §58), not to downgrade the data model.
ADR-0003 · The as_of repository layer is the only data access path
Decision. All reads of L1–L5 go through repository methods that require an as_of
timestamp. Raw SQL outside src/elab/repo/ is forbidden and tested against.
Alternatives. Convention plus code review; a separate "historical" query module.
Why. Leakage is the dominant failure mode of this class of system (FM-01 … FM-06, FM-12 in
BACKTESTING's taxonomy). Two code paths guarantee eventual divergence. One path, with
production as the special case as_of = now(), makes the backtest and the live system the same
program.
Falsifier. A demonstrated performance problem that cannot be solved within the repository abstraction. Profiling required first, per build spec §70.
ADR-0004 · Inference engine chosen by benchmark, not preference
Decision. PyMC + nutpie is the leading candidate; CmdStan via cmdstanpy is the fallback.
The choice is made in Phase 1 by benchmarking the same model in both on this hardware.
Criteria. Effective samples per wall-clock second on the Phase 1 model; divergence behaviour on the hierarchical funnel; peak RSS under 16 threads; ease of the non-centred reparameterisation; quality of diagnostics.
Why. No GPU, so the JAX/GPU path is irrelevant here. Between the CPU options the difference is empirical, and the build spec explicitly forbids choosing a framework because it is fashionable.
Falsifier. The benchmark. It is a committed artefact, re-runnable on other hardware.
Outcome, 2026-09-23 — PyMC's own NUTS, not nutpie
Measured on Skynet10, Phase 1 model, 3 seeds, 4 chains × 1000 draws
(experiments/adr0004_sampler_benchmark.json):
| Sampler | median ESS/s | median seconds | peak RSS |
|---|---|---|---|
| pymc | 180 | 7.6 | ~450 MB |
| nutpie | 93 | 9.1 | ~650 MB |
This contradicts what this ADR expected. nutpie was written down as "the leading candidate" on reputation, and on this model it is roughly half the throughput at ~45% more memory. The ADR committed to deciding by benchmark rather than by fashion; the benchmark decided against the favourite, so the favourite loses.
Scope of the result, honestly. One model, three seeds, a small graph (~250 latent days, 5 pollsters). nutpie's advantage is generally reported to grow with model size, and the Phase 5 Senate model is much larger. This decision is therefore provisional and scoped to Phase 1–4, and the benchmark is re-run before Phase 5 rather than inherited. Re-running it is one command, which is the point of committing it.
Re-run on the Phase 5 model, 2026-09-27: the decision stands, for a better reason
The provisional status is lifted. scripts/benchmark_samplers.py --real-race benchmarks the most
heavily polled real race in the corpus — US-PA 2024, 134 polls, 64 firms — which is the "much
larger" model this ADR was waiting for. Three seeds each
(experiments/adr0004_sampler_benchmark_phase5.json, host-tagged):
| sampler | seconds | min ESS | ESS/s | max R-hat | peak RSS |
|---|---|---|---|---|---|
| pymc | 23.8–24.6 | 1,239 | 50.3–52.0 | 1.0062 | 426–453 MB |
| nutpie | 8.9–14.8 | 451 | 30.4–50.6 | 1.0143 | 572–677 MB |
nutpie's advantage did grow with model size, in wall clock: it finishes in roughly a third of the
time. But it draws about a third of the effective sample, so throughput per unit of information is
a tie at about 50 ESS/s — and its maximum R-hat of 1.0143 exceeds this project's own 1.01
diagnostic gate, on every seed. A sampler whose output passes_gates would reject is not faster; it
is a run that has to be repeated.
So the original conclusion holds and the reason is now the stronger one: on the small model nutpie lost
on throughput, and on the large one it loses on mixing. Memory is also 40–50% higher, against a
worker_memory_mb of 2,048 and four concurrent workers.
And the script had been broken since 0.5.0. fit() gained a required start_prior when ADR-0007's
fundamentals prior moved onto the race intercept (FM-84), and this benchmark was never updated — it had
been raising TypeError on every invocation, unnoticed, because nothing runs it. That is the second
script found in this state on the same day (FM-92 was the first), from the same signature change.
CmdStan remains unbenchmarked: it requires compiling a toolchain and is recorded as an omission rather than silently treated as compared.
ADR-0005 · Compositional (ALR) scale, with a Dirichlet-multinomial alternative tested
Decision. Model latent support in additive-log-ratio space with undecided/other as the reference category.
Alternatives. Raw share with Gaussian noise (rejected: impossible intervals near the boundary, and it mishandles lopsided races); direct margin modelling (rejected: throws away the undecided category, which the model needs to reason about explicitly); Dirichlet-multinomial on the simplex (retained as a tested competitor).
Why. Vote shares are compositional. Multi-candidate races and third-party dynamics are first-class requirements, and ALR handles them while reducing to a plain logit in the two-candidate case.
Falsifier. Out-of-sample comparison in Phase 4.
ADR-0006 · No workflow orchestrator before Phase 7
Decision. Plain Python entry points and make targets through Phases 1–6. Revisit
Airflow/Prefect/Dagster at Phase 7 when scheduled research automation actually exists.
Alternatives. Adopt an orchestrator immediately.
Why. On a single 16-thread machine with no running DAGs, an orchestrator is operational overhead with no benefit, and it would be adopted before the shape of the workload is known.
Falsifier. Phase 7's scheduling requirements, or any point at which job dependency management becomes the dominant source of bugs.
ADR-0007 · Fundamentals are a prior, not a weighted-average ingredient
Decision. The fundamentals model supplies the prior on the race intercept α_r; polls
update it through the likelihood. No hand-tuned fundamentals-to-polls blending schedule.
Alternatives. The common public-model approach of an explicit time-varying weight between a "fundamentals forecast" and a "polling forecast".
Why. The explicit blend requires inventing a decay schedule that has no principled estimand and that becomes a tuning knob — and tuning knobs are where overfitting enters. A prior-plus-likelihood formulation makes the weight fall out of the arithmetic: it declines automatically as polling information accumulates, at a rate implied by the relative precisions.
Falsifier. If out-of-sample evaluation shows an explicit blend beats the Bayesian update — which would indicate the likelihood's variance is misspecified — investigate the misspecification rather than adopt the knob.
ADR-0008 · Compute is configuration, not code
Decision. Worker counts, memory caps, chain counts, draw counts, and paths live in
config/compute.yaml. Benchmarks are tagged with host identity.
Why. This machine is a 16-thread laptop with 30 GB RAM and exhausted swap; earlier project notes assumed a dual-EPYC host with 1 TB. Both should run the code unchanged. Untagged benchmarks compared across hosts would produce false performance conclusions.
Falsifier. None expected; this is a hygiene decision.
ADR-0009 · Primary pollster releases are the authoritative poll source
Decision. The polling database is built primarily from pollsters' own releases, with archives and aggregators used as discovery and cross-check layers, never as terminal sources.
Alternatives. Build on an aggregator feed (the historical default).
Why. Verified on 2026-09-22: the FiveThirtyEight raw poll CSV endpoint no longer serves CSV (DATA_SOURCES §1). Beyond availability, aggregator rows lack the methodology fields the model requires, and they embed the aggregator's own entity-resolution and inclusion decisions — which would silently import another organisation's judgement into ours and make independent evaluation partly circular.
Cost. This is the largest work item in the project and the main justification for the LLM extraction pipeline.
Falsifier. If a licence-clean, methodology-complete archive with verifiable provenance
becomes available, it supersedes the discovery layer — but primary documents remain the
resolution target for any row marked verified.
ADR-0010 · Holdout is a metered, exhaustible resource
Decision. A holdout ledger tracks evaluations per cycle; the most recent complete cycle starts sealed; cycles evaluated by more than 20 specifications are marked burned and may not be cited as out-of-sample evidence.
Alternatives. Trust the analyst; rely on cross-validation alone.
Why. An automated challenger generator will otherwise convert ten historical elections into a training set by slow attrition, and every metric will keep improving while real accuracy decays (FM-31, FM-32). The threshold of 20 is a policy choice, not an estimate — it is deliberately conservative and is configuration, not a constant in code.
Falsifier. A defensible multiplicity-corrected procedure that permits more evaluations without inflating false discovery would replace the crude counter.
ADR-0011 · LLMs may not originate numbers, enforced by database grants
Decision. Any LLM-facing process runs as elab_narrative, which has no INSERT privilege on
any data or forecast table.
Alternatives. Policy and code review.
Why. The build spec's most important architectural constraint deserves an enforcement mechanism that survives a bug, a refactor, or a prompt injection. A permission denial is such a mechanism; a convention is not.
Falsifier. None. If this constraint is ever relaxed, the project has changed into something else.
ADR-0012 · Phase 1 races: 2022 Pennsylvania Senate, plus 2022 Utah Senate as the awkward case
Decision. The vertical slice is built on the 2022 Pennsylvania Senate race (Fetterman–Oz), and exits only when it also works for the 2022 Utah Senate race (Lee–McMullin).
Selection criteria (from PHASE1_PLAN M1.5): dense public polling, a certified result, no mid-cycle candidate replacement, sources whose licence is already verified.
Why Pennsylvania 2022. Verified on 2026-09-22: its Wikipedia article carries 7 polling tables, and the poll rows link directly to primary pollster documents. It is a conventional two-major-party race with a certified result and no candidate substitution, so the easy path through every layer is genuinely exercisable. Fetterman's stroke is a campaign shock rather than a data-model problem — he remained the candidate — which makes it a useful late-movement case without breaking candidate identity.
Why Utah 2022 is the required second race. A pipeline validated only on the easy case is not validated. Utah 2022 had no Democratic candidate: the state Democratic party endorsed Evan McMullin, an independent. That single fact breaks assumptions that a two-party design absorbs silently:
- there is no two-party margin, so a naive
D% − R%is meaningless and the "uncontested races produce missing, not zero" rule (STATISTICAL_MODEL §8, FM-17) has to actually fire; - the compositional parameterisation must handle a major candidate with no party of the usual kind, exercising ADR-0005 rather than degenerating to a logit;
- the partisan baseline for Utah is uninformative about a Lee-versus-independent contest, so the fundamentals prior has to express that rather than confidently predicting a Republican blowout margin against a Democrat who does not exist.
It also has 9 polling tables, so it is awkward without being data-starved — the difficulty is structural, not merely thin.
What this deliberately does not test. Sparse polling (near-zero polls) and mid-cycle candidate replacement remain untested after Phase 1. They are Phase 5/6 concerns, and are recorded here so their absence is a known gap rather than an assumed pass.
Falsifier. If Utah 2022's poll records turn out to be too sparse or too poorly documented to exercise the multi-candidate path, substitute another race with a non-major-party contender and record the substitution here.
ADR-0013 · Remote compute is a calculator, not a client
Decision. Heavy model fitting runs on skynet-three (dual EPYC 7742, 256 threads, 1 TB
RAM) via a standalone worker that receives already-filtered observation lists and has no
database access.
Alternatives. (a) Fit locally — 671 fits would take hours on 16 threads instead of three minutes. (b) Replicate the database on the remote host and let the worker query it directly.
Why not (b). A worker that can query the database is a second read path. ADR-0003's whole
argument is that one as_of-parameterised path is what makes leakage structurally impossible;
a second path re-opens exactly the hole that was closed, and on a machine where the leakage
tests do not run. Filtering happens locally through the repository, and what crosses the
network is numbers the worker cannot add to.
The gate travels with the jobs. export_fit_jobs.py writes DIAGNOSTIC_GATES into the
job file and the worker reads it. A hardcoded copy would drift from the local gate the first
time either changed, and the drift would be invisible — both paths would keep producing
plausible numbers under different standards. tests/unit/test_remote_gate_parity.py asserts
the worker holds no literal threshold of its own.
Politeness is a requirement, not a courtesy. The host runs four RTX 3090s serving
llama-server inference whose latency users notice. The worker never touches a GPU (PyMC's
default NUTS is CPU-only; no JAX path), runs at nice 15, pins BLAS to one thread per worker,
and uses 40 workers (160 threads) leaving ~96 free. Measured during a full run: GPU
utilisation and llama-server CPU share unchanged.
Falsifier. If a future model needs data the exporter cannot anticipate, the answer is to extend the export, not to give the worker a database handle.
ADR-0014 · Cycle-level bootstrap for every model comparison
Decision. Paired model-versus-baseline comparisons resample cycles, not races.
Why. Races within a cycle share the national polling error, so they are not independent. Resampling races would treat ~135 correlated observations as 135 independent ones and shrink the confidence interval by roughly the square root of the average cycle size — turning noise into significance.
What it cost. The reported CI on the log-loss difference is [−0.0395, −0.0087]. Race-level resampling would have produced a far tighter interval around the same estimate and an unjustified claim of precision.
Falsifier. A demonstration that within-cycle correlation is negligible for a given metric, which would have to be measured rather than assumed.
ADR-0015 · A district's crosswalk is derived from the source that needs it, not licensed
Decision. The precinct-to-district mapping used to compute a district's partisan baseline is derived from the same precinct file that supplies the votes — a precinct's district is stated on its own US HOUSE rows — rather than obtained from a published crosswalk.
Why. The widely used crosswalk is the Redistricting Data Hub's, and source_registry carries it
with enabled = false: its terms are a standing restriction on what this project may do with its own
outputs, which is a cost paid forever for a convenience used once. The alternative was assumed to be
"no district baselines", which is why boundary-crossing priors were handled by widening for so long.
It was not: a precinct's district and its presidential votes are in the same file, keyed the same
way, so the join is internal. New York matches 12,661 of 12,862 precincts.
What it cost. Every state's file has to be trusted, and several cannot be. Four checks were
needed before the output was usable, each catching what the previous ones could not: the precinct
match rate (Oklahoma, Rhode Island), agreement with the state's certified presidential margin
(Washington matched 100% of precincts and was fourteen points out), vote coverage with both a floor
and a ceiling (Louisiana was right by luck on half the state; Indiana came to 105% because its
file reports some votes twice, which no margin check can see), and crosswalk unambiguity (Ohio lists
every precinct under all fifteen districts). Coverage is 358 of 435 districts across 40 states, and
each of the ten exclusions is named in experiments/district_baseline.md.
That is a worse coverage figure than a licensed crosswalk would give and a better one than the project had, which was zero. The measured effect on the districts it covers is 3.17 points of mean absolute error against 13.11 for the prior it replaces.
Falsifier. A permissively licensed crosswalk — a state publishing precinct shapes with district assignments, or a Census block-assignment path that does not need precinct geometry — would make this reasoning unnecessary for the states it covers. The mechanism here would stay as the fallback, because it needs nothing beyond the returns themselves.
ADR-0016 · CRPS is the primary metric for a challenger that changes a distribution's shape
Decision. A challenger whose mechanism changes the width or shape of a predictive distribution, rather than its central estimate, is evaluated on CRPS as its pre-registered primary metric. Log loss remains primary for a challenger that changes which outcome is favoured.
Why, from what the two scores measure. Log loss reads one number out of the distribution: the probability assigned to the outcome that happened. On a race whose winner is not in doubt that number is near 1 whatever the margin distribution looks like, so log loss is near zero however wrong the margin is and however narrow the interval. It cannot distinguish an honest interval from a dishonest one when the sign is right, and an interval is exactly what a calibration mechanism changes.
CRPS integrates the squared difference between the predictive CDF and the outcome's step function. It is a proper score, it is in the units of the quantity being forecast, it rewards a distribution that is both well-centred and honestly wide, and it reduces to absolute error for a point forecast — so a model cannot improve it by being vague.
What it cost. c9_sparse_poll_variance was pre-registered on log loss, improved coverage from
0.33 to 0.67 on one fold and CRPS on all three, and was rejected because log loss got worse
everywhere (FM-72). This ADR does not rescue it: a metric adopted after seeing a result is not
evidence about that result, and the mechanism can be re-registered only against a cycle not yet used
for it. The cost of that discipline is a rejected fix to a defect this project has measured, which is
the price of the rule being worth anything.
Falsifier. A demonstration that CRPS and log loss rank distributional challengers the same way in
this domain would make the distinction unnecessary. c9 is a counterexample in the other direction,
which is why this exists.
ADR-0017 · A defect is fixed; only a hypothesis faces the gate
Decision. A change justified by a stated mechanism plus a measurement of a defect is a fix: it is versioned, documented, republished, and does not go through the promotion gate or consume holdout. A change justified by winning a comparison is a challenger and faces the gate. Which one a change is depends on its justification, not on how much of the model it touches.
Why. The gate exists to stop one thing: searching over plausible variants until one wins by luck, on the only test data that will ever exist (ADR-0010, FM-31, FM-32). That is a guard on model selection. It is not a licence to keep publishing a number that is wrong on its own terms.
The project has now been on both sides of this line in a single day, correctly once and wrongly once.
Correctly: effective_n was computing neither of the two things its name could mean, and the wrong
formula was replaced without a gate because a wrong formula is not a hypothesis (FM-73). Wrongly: the
variance budget was missing the larger of its two polling-error terms, a term the same codebase had
been estimating all along and the chamber simulator had been using — and instead of restoring it,
three rounds of c8_bias_centre and one of c9_sparse_poll_variance argued about the symptom through
the gate, were refused on a metric that could not see what they addressed, and left a nominal 90%
interval covering 68% for three more days (FM-79).
The test. Ask what would have to be true for the change to be wrong.
- "The estimator is computing something other than what it claims" — a fix. Checkable by reading it.
- "A term the model needs is absent from the budget" — a fix. Checkable against the decomposition that defines the budget.
- "Polling error is conditional on how uncompetitive a race is, and that will hold next cycle" — a hypothesis. It can only be checked against outcomes, and against outcomes it has not yet seen.
A useful corollary: a fix must be stateable without reference to any score. If the only argument for a change is that a metric improved, it is a challenger no matter how obviously right it looks.
What a fix still owes. Everything except the gate. A semver bump, because it changes what the
version promises (FM-69). A refit of every affected cycle, walk-forward. Republication. A
FAILURE_MODES entry with the numbers before and after. And a measurement of the effect, reported
whether or not it flatters the change — sigma_race took 90% coverage from 0.680 to 0.895 and left
the +3.51-point lean exactly where it was, and both halves of that are in the record.
Falsifier. A change that satisfies the fix test and makes the model demonstrably worse out of sample. That would show the distinction is not load-bearing and everything belongs in the gate.
ADR-0018 · The gate has a calibration floor, and a threshold per metric
Decision. Two additions to elab.registry.promotion:
- A calibration floor. A challenger that supplies realised coverage missing any nominal band by
more than
COVERAGE_TOLERANCE(0.05) cannot be promoted, whatever the comparison says. Coverage must be supplied for anything scored against real outcomes (require_coverage, default true); the synthetic studies that calibrate the gate's own false-discovery rate waive it, because a null challenger is a vector of score differences with no predictive distribution behind it. MIN_EFFECT_BY_METRIC. The effect threshold is looked up from the challenger's own metric — 0.005 for log loss, 0.073 for CRPS — and a metric with no entry raisesUnanchoredMetricrather than borrowing one.
Why the floor. The gate only ever compared specifications. Nothing in it could notice that the champion's nominal 90% interval covered 0.680 and its nominal 50% covered 0.259, because no challenger had been offered that did better — and the one change that would have fixed it makes log loss worse on three cycles of four, so it would have been refused (FM-79). A gate that can reject the repair of a defect it cannot see is not controlling quality, it is controlling change.
Only under-coverage blocks. An interval that covers more often than it claims is conservative; that costs sharpness, CRPS charges for it, and it is not a claim that turns out to be false. A floor that refused it would refuse honesty in the one direction this project has never erred in.
Why the per-metric threshold. 0.005 is in log-loss units. Applied to CRPS, which is in points of
margin, it is roughly a fifteenth of the smallest difference worth noticing — every challenger passes,
and the gate silently stops controlling effect size at the moment ADR-0016 changes the metric. Both
thresholds are now about a fifth of the champion's measured edge over the same opponent, a recency-
and size-weighted polling average widened by empirical error: 0.027 → 0.005 for log loss, and
0.364 → 0.073 for CRPS, measured on 246 races with the champion ahead on 4 of 4 cycles
(experiments/crps_edge.json, scripts/measure_crps_edge.py). Refusing an unanchored metric is the
part that matters: it forces the measurement before the first evaluation rather than after it.
Where the current champion stands. single_race_latent/0.4.0 covers 0.439 / 0.768 / 0.886 against
0.50 / 0.80 / 0.90. The 90% and 80% bands are inside tolerance; the 50% band is not, missing by
0.061. That is where a location error shows up first — the middle of the distribution is the most
sensitive to being off-centre — and the +3.53-point lean is untouched. The floor's verdict and the
diagnosis agree, which is the check on both, and it is asserted as a test so it cannot drift silently.
Falsifier. A promotion the floor blocks that later turns out to have been right — a specification whose measured under-coverage on four cycles was sampling noise. With four effective observations the tolerance is a policy and not an estimate, and it is set loose enough to say so.