Election LabU.S. forecasts, with every number traced to its run

Methodology

the documents the model was built against, and what the numbers mean

The documents

The repository's own methodology, rendered from disk at request time rather than summarised, so what you read is what the code was built against: the statistical specification, a numbered log of every defect ever found in this system and what it cost, the comparison with the other forecasters, and a table of every fitted parameter the live forecast is standing on.

What are the two numbers?

A NOWCAST answers 'if ballots were cast today'. An ELECTION-DAY forecast answers 'what happens in November'. They differ by the movement still to come, so the second is the wider of the two wherever there is time left. They are never averaged, never combined, and never shown as one with a caveat. Where only one exists, the race or chamber is withheld rather than shown half-filled.

Where does a probability come from?

A correlated Monte Carlo simulation over a fitted model, not from a formula applied to a margin. Races move together, so the simulation draws a shared national error alongside each race's own; treating races as independent would make a 435-seat chamber look knowable to within a couple of seats.

What is the uncertainty made of?

Latent opinion from the fit; correlated industry-wide polling error; turnout composition; late-breaking undecideds; between-specification uncertainty; and, for an election-day forecast, drift between now and the election. Each is stored with the forecast, so a probability whose dominant input is an unfitted placeholder is visibly that rather than quietly that.

Does a language model write any of these numbers?

No. An LLM never originates a forecast, a probability, or any number on these screens. Every figure traces to a simulation run with a recorded seed, model version and commit. Language is used for description and for proposing model changes, and a proposed change is only adopted if it passes the same pre-registered statistical gate as any other.

How does a new model get adopted?

Against a champion, on cycle-level bootstrap with Benjamini-Hochberg control of the false discovery rate, requiring replication across cycles, a stated mechanism for why it should work, and a minimum effect size. The criteria are registered before the comparison is run. A challenger that wins on a metric chosen after seeing the data has not won -- and the metric depends on what the change does: log loss where a mechanism changes which side is favoured, CRPS where it changes how wide the distribution is, because a score that reads only the winner cannot see whether the interval around it is honest. Four challengers have been evaluated and four rejected, including two whose mechanism was measured and whose effect went the wrong way on the registered metric. The refusals are recorded with the same weight as a promotion would be.

How do you know the backtest is honest?

Walk-forward, with every input resolved as of the forecast date through a bitemporal repository layer: a 2014 forecast cannot see a poll recorded in 2015 or a candidate's later party. Oracle canaries are planted to confirm the leak detection fires. The holdout is treated as a consumable resource with a recorded budget, because a test set you look at twenty times is a training set.

What is this system unable to tell you?

Whether the polling industry is wrong in a direction all of it shares -- that is estimated from past results as a variance term, not detected in the current polls. On that, be concrete: scored against certified results, this model's forecasts have leaned Democratic by about four points in every cycle measured, and three pre-registered attempts to correct it have been refused because nothing estimable before a cycle predicts that cycle's lean. The bias is documented rather than fixed, which is why an interval here is wide. It also cannot tell you about a race it has too little evidence for: with no polls at all it refuses to publish a race forecast rather than producing a number without evidence behind it.

Can a forecast be reproduced?

Every run records its model version, git commit, random seed and configuration hash. Raw sources are archived content-addressed and immutable, so the inputs to an old forecast are the bytes that were actually fetched rather than whatever the source serves today.