FutureBallotU.S. election forecasts, with every number traced to its source

Data sources

Every source, its licence, what is taken from it, and the ones that are deliberately not used with the reason.

DATA SOURCES

All reachability and licence statements below were probed on 2026-09-22 from this machine and are recorded with what was actually observed. Nothing here is assumed. A source is not fetched by the pipeline until its row in source_registry has enabled = true, which requires a human licence decision (DATABASE_SCHEMA §2).


1. The finding that changes the plan

The FiveThirtyEight raw poll feed is gone.

GET https://projects.fivethirtyeight.com/polls/data/president_polls.csv
  -> HTTP 200, Content-Type: text/html
  -> body is an ABC News HTML page, not CSV

The endpoint that nearly every open-source election model in the last decade was built on now returns a marketing page. The GitHub archive fivethirtyeight/data is still live and still CC BY 4.0 + MIT (verified by reading its README), but its polls/ directory contains only:

File Size Coverage
pres_pollaverages_1968-2016.csv 50.4 MB presidential polling averages, not poll records
pres_primary_avgs_1980-2016.csv 30.1 MB primary averages
2024-averages/ small 2024 averages

Last commit touching polls/: 2024-09-13. The pollster-ratings/ directory is present and is the most useful surviving artefact there.

Consequence

The polling database cannot be built by downloading one CSV. It must be assembled, and the assembly is the largest single work item in the project. Ranked by defensibility:

  1. Primary pollster releases — the authoritative source, and the only one that supplies the methodology fields in build spec §5 (mode, weighting, LV screen, question wording). Highest fidelity, highest labour. This is where the LLM extraction agents earn their keep.
  2. Surviving third-party archives with a clear licence. Each must be verified for provenance; a mirror is not an authority, and a fork of 538's CSVs is evidence of what 538 published, not evidence of what the pollster published.
  3. Academic archives (Roper iPoll, ICPSR) — high quality, licence-restricted, need the user's authorisation.
  4. Wikipedia election poll tables — broad coverage, CC BY-SA, crowd-maintained. Usable only as a discovery layer that points at primary documents, never as a terminal source. Every Wikipedia-sourced row must resolve to a primary artefact before it is verified.

This ordering is itself a design decision (ADR-0009) and it is why the architecture treats extraction-from-primary-documents as a first-class pipeline rather than an afterthought.


2. Verified-reachable sources

Source Probe result Licence status Use
FEC OpenFEC API api.open.fec.gov 200 JSON with DEMO_KEY; key installed (§2.2) US Government work, public domain Candidate registration, fundraising, receipts, cash on hand, outside spending (build spec §50)
Census API api.census.gov 200 Public domain ACS demographics for correlation factors (§10 of STATISTICAL_MODEL)
Data.gov catalog API catalog.data.gov 200 JSON, no authentication Catalog metadata is public; per-dataset licences are not settled by the catalog (see §2.1) Source discovery — finds federal, state, and county datasets, including precinct-level results
Census TIGER www2.census.gov/geo/tiger/TIGER2024/CD/ 200 directory listing Public domain Congressional district shapefiles → boundary_hash, redistricting lineage
Harvard Dataverse API dataverse.harvard.edu/api/search 200 JSON Per-dataset; MIT Election Data & Science Lab datasets are typically CC0 — verify per DOI Certified returns, county and precinct level
MIT Election Lab electionlab.mit.edu 200 see above Landing page for the Dataverse DOIs
Congress.gov API api.congress.gov 403 API_KEY_MISSING without a key; 200 with DEMO_KEY; key installed (§2.2) Public domain data, key-gated Member/incumbency history
State election offices (e.g. electionreturns.pa.gov) 200 Varies by state; generally public record Authoritative certified results; the ground truth for election_result
fivethirtyeight/data GitHub 200 CC BY 4.0 (data) / MIT (code), confirmed in README Pollster-ratings history, historical averages, benchmark series
Redistricting Data Hub 200 Registration + terms acceptance required Precinct-level results for boundary disaggregation (§8)
Roper Center iPoll 200 Subscription. Requires institutional access Deep historical poll archive with methodology
Wikipedia REST API 200 JSON CC BY-SA 4.0 Discovery layer only (see §1)

2.1 Data.gov catalog API — discovery, verified 2026-09-22

No authentication. Documentation at resources.data.gov/catalog-api. Endpoints confirmed by probe:

Endpoint Observed
GET https://catalog.data.gov/search?q=…&per_page=… 200 JSON: {after, results[], sort}; each result carries a DCAT record under dcat
GET https://catalog.data.gov/api/keywords?size=…&min_count=… 200 JSON: {keywords:[{keyword,count}], total}
GET https://catalog.data.gov/harvest_record/<uuid4> (+ /raw, /transformed) Per documentation; /transformed is the recommended source of current DCAT

Other parameters documented: org_id, org_type, keyword, after, spatial_filter, sort. Pagination is cursor-based on the after token, not offset-based — the ingest client must follow the cursor rather than compute page numbers. The legacy CKAN path /api/3/action/package_search returns 404, so this is a distinct API and CKAN client libraries will not work against it.

Why this matters more than it first appears. A query for election results returned, in the first eight hits, precinct-level results from King County (WA) and statewide results from Connecticut, alongside the NIST Election Results Reporting XML Schema. Precinct-level results are the scarce input for redistricting disaggregation (§8 of STATISTICAL_MODEL, FM-16) — the thing that makes district baselines comparable across boundary changes. A no-auth, queryable index of state and county open-data portals is therefore a genuine discovery channel for the project's hardest data problem, not merely a directory of federal CSVs.

Role. This is the backbone of the Source Discovery Agent (build spec §25). It is a discovery layer: it returns metadata and pointers, and the pointed-at data is then fetched, archived, and verified through the normal L0 path.

Three limits, stated plainly.

  1. It does not settle licensing. Most returned DCAT records carried an empty license field. Discovery via the catalog does not flip source_registry.enabled to true; a licence decision still has to be recorded per dataset. The catalog tells us a dataset exists, not that we may use it.
  2. Coverage is uneven and metadata quality varies. Many records point at municipal Socrata portals with inconsistent schemas and no methodology documentation. Discovery volume is not data quality.
  3. It does not replace the API keys. catalog.data.gov (the dataset catalogue) and api.data.gov (the shared federal API gateway that issues keys for OpenFEC) are different services. The FEC and Congress.gov keys are still needed for those APIs at volume.

2.2 API credentials — resolved 2026-09-22

One api.data.gov key serves both Congress.gov and OpenFEC. Both sit behind the api.data.gov umbrella gateway (visible in the x-api-umbrella-request-id response header), and the same key authenticates against both, via the X-Api-Key header. The user's existing key was located and installed as ELAB_CONGRESS_GOV_API_KEY and ELAB_DATA_GOV_API_KEY in the git-ignored .env.

Verified live:

Service Auth Observed x-ratelimit-limit
api.congress.gov/v3 X-Api-Key header 20000
api.open.fec.gov/v1 same key, same header 60
either, with DEMO_KEY — 10

The two services rate-limit the same key by more than two orders of magnitude, and the header does not state its window. A limiter that assumed "requests per hour" from one service would be wrong for the other by a wide margin in whichever direction the window differs. This is direct vindication of the fetcher design decision in PHASE1_PLAN M1.4: read x-ratelimit-limit / x-ratelimit-remaining from each response and adapt per service, rather than hardcode a rate per source. The source_registry.rate_limit_rps column becomes a conservative ceiling, not the operative control.

Still needs a human before it can proceed

Item Why Ask
Census API key Not yet obtained. ACS works unauthenticated at low volume; a key is needed for bulk pulls Free self-service key
Roper iPoll access Subscription Does the user have institutional access?
Redistricting Data Hub account Terms acceptance is a legal act User must accept
CES / ANES cces.gov.harvard.edu returned 403 to a bare curl; may be UA-filtered rather than unavailable Re-probe with a proper UA; registration likely

Build spec §58 lists "credentials required" and "licensed data require authorization" as legitimate reasons to interrupt. These are those.


3. Sources deliberately not used

Source Reason
RealClearPolitics Aggregated content; scraping is against site terms and a lawful alternative (primary releases) exists
Cook Political Report / Sabato's Crystal Ball ratings Proprietary and paywalled. Also circular as features (build spec §40): expert raters consume polls and public forecasts, so their ratings leak the very information we are trying to evaluate independently
Any current competitor's live model output Benchmark only, and only where published openly with a usable licence. Never an input
Scraped paywalled newsletters Prohibited

4. Prediction markets

Status: not confirmed. predictit.org/api/marketdata/all/ returned 200 JSON. Two Kalshi and Polymarket hostnames I guessed failed DNS resolution — that is evidence I guessed wrong hostnames, not evidence the APIs are unavailable, and the docs must be read before any claim is made about them.

Regardless of availability, markets are not ingested as a model input until they pass the test in build spec §41: do they add out-of-sample predictive value beyond polls and fundamentals? Markets partly price public forecasts, so naive inclusion creates circularity (FM-23). They enter the experiment registry as a hypothesis, and the null result is an acceptable outcome.


5. Fundamentals series

Series Source Vintage handling
Unemployment, payrolls, CPI, real disposable income, GDP BLS / BEA / FRED Vintage-correct series mandatory (ALFRED-style real-time vintages where available). Using revised history is look-ahead bias, FM-03
Presidential approval Aggregated from primary pollster releases Same poll pipeline; approval is just another question type
Fundraising FEC OpenFEC Filing-date-aware: a Q3 report is not knowable before its filing date
Special-election results State offices build spec §49 hypothesis
Registration statistics State offices (coverage is uneven; ~20 states report party registration) Coverage gaps documented, never imputed silently

The FEC rule is worth stating twice: a campaign finance report becomes knowable on its filing date, not on the last day of the period it covers. The repository enforces this.


6. Coverage gaps, stated honestly

No gap is filled by fabrication. A missing value stays missing, and its absence propagates into the uncertainty.


7. Fetching discipline

  1. Read robots.txt and terms before first fetch; record robots_checked_at.
  2. Prefer official APIs and bulk files over HTML scraping.
  3. Honour rate limits from source_registry.rate_limit_rps; exponential backoff on 429/5xx.
  4. Identify the client honestly in the User-Agent, with contact information.
  5. Archive raw bytes before parsing, always (DATABASE_SCHEMA §2). Parsing is re-runnable; fetching is not.
  6. Cache by content hash — re-fetching identical bytes is a no-op.
  7. All fetched content is untrusted input. See SECURITY.md.

FEC biennial report: the workbook, not the PDF

The FEC publishes Federal Elections as both a PDF and a workbook. This project reads the workbook. The two carry the same data, but the PDF carries it as typography: the earlier parser needed a tuned column offset, a wrapped-row joiner and a regular expression for candidate names, and each of those is a guess whose failure mode is a plausible number rather than an error. In the workbook, fusion totals, runoff rounds and ranked-choice rounds are each a labelled column, so the only question left is which column the law says decides the election.

Document ids, verified against the FEC's own per-cycle index pages:

Cycle Workbook id File
2014 1700 federalelections2014.xls
2016 1890 federalelections2016.xlsx
2018 2706 federalelections2018.xlsx
2020 4228 federalelections2020.xlsx
2022 5676 federalelections2022.xlsx

2024 is not listed because it is not published. Both URL patterns were checked on 2026-09-23: /election-results-and-voting-information/federal-elections-2024/ returns 404 and /election-and-voting-information/federal-elections-2024/ returns 200 with no document links. That is a checked absence, not an assumed one.

Do not guess document ids. An earlier attempt brute-forced roughly 1,400 of them, which got this client rate-limited for half an hour in violation of section 7 above, and produced five wrong ids that looked right: four served MP3 recordings of open meetings, and the fifth served a valid PDF of a 2003 press release about campaign spending, filed under the id guessed for 2014. The content-type header did not catch the last one, because it really was a PDF. The fix is in two parts: Fetcher(expect_magic=...) checks the body rather than the header, and ids come from the index pages.

What the report settles

Against the Clerk of the House's official statistics, the parser reproduces the House split exactly for all five cycles -- 435 seats, correct D/R division, no third-party or unresolved seats. Rules that each fix a real contest:

Governors: three sources, none of them federal

The FEC's remit stops at Congress, so there is no single authority for governors and every one of this project's 1,180 governor results rested on one Wikipedia pipeline. Three independent sources were found; between them they certify 561 results (47.5%), up from none.

Source Covers Access Form
MEDSL state-office returns 2022, 2024 open (CC0) precinct-level, 0.9–1.3 GB per cycle
MEDSL state-office returns 2016, 2018, 2020 guestbook same
OpenElections 2004–2024, 46 states unevenly open county, precinct or statewide CSV
MEDSL "US Governors 1775–2020" winners only guestbook who held office each year

MEDSL mirrors the same returns on GitHub, without the guestbook. Three of the five Dataverse state-returns files sit behind guestbookID 458; the GitHub organisation publishes the same data openly, and checking it turned two of those three into ordinary downloads:

Cycle Open source File
2016 MEDSL/state-returns stateoffices2016.csv, 2 MB, already totalled to state level
2018 MEDSL/2018-elections-official STATE/STATE_precinct_general.zip, 78 MB zipped
2020 none — MEDSL/2020-elections-official is deprecated and redirects to Dataverse

So the guestbook blocks exactly one cycle, 2020, not three. Finding that took one call to the GitHub API, after I had already written the gap up as a decision for the user to make. It is the fourth instance this project has recorded of a gap that was really an unsearched source (FM-43), and the first where I had gone as far as asking someone else to act on it.

Where it does apply, the guestbook is not worked around. It is a condition the depositor set and a representation made to a named institution. The datasets themselves carry CC0 1.0 and no terms of use at all, so the question is one of access rather than licensing.

The last of these is not usable as it stands. It lists who held office in each year rather than who won each election, so it includes appointed and successor governors — Jeff Colyer appears for Kansas in 2018 and 2019 without ever winning an election. Comparing it to election results would manufacture disagreements wherever a governor resigned.

Two guards decide what is usable

Neither source is trusted on its own, because both are compiled by summing smaller units and summing is where a plausible wrong number comes from.

What each source needed

MEDSL codes a fusion candidate's minor line as "OTHER" under party_simplified, so Richard Blumenthal appears twice in Connecticut and summing only the first put his share 0.8 points low. OpenElections needed three fixes, each found by a state that looked empty or impossible: several states elect a joint ticket and label the office "Governor and Lieutenant Governor", so excluding anything mentioning a lieutenant governor dropped four states entirely; Massachusetts's "Governor's Council" is a different body and matched on the bare word; and states spell parties their own way, "DEMOCRATIC PARTY" in New Mexico against a bare "D" elsewhere, where an unrecognised spelling silently drops a candidate from the two-party vote.

The file for each state and year comes from data/openelections/file_index.json, built by listing each repository's tree once rather than by guessing paths. A guessed path returns 404, and a 404 read as "no data" is how a gap gets recorded that is not there — which this project has done often enough to name it, as FM-43.

Generic ballot after 2016

The 538 trendline this project uses (generic_topline_historical.csv) ends 2016-11-06. Without a staleness limit that produced a confident zero swing for every later cycle, which is FM-45.

Searched, and not found in a current source:

candidate result
fivethirtyeight/data, congress-generic-ballot/ one file, ends 2016
fivethirtyeight/election-results actual House results, not polls
projects.fivethirtyeight.com/polls*/data/generic_ballot_polls.csv returns HTML; the service is gone
Wikipedia House-elections pages, current revisions no generic-ballot section
Wikipedia House-elections pages, election-eve revisions none either — it was never on that page
Wikipedia "YYYY United States elections", current and election-eve only 2026 carries one

Found in the Internet Archive. The 538 endpoint was captured while it was live: web.archive.org/web/<ts>id_/https://projects.fivethirtyeight.com/polls/data/generic_ballot_polls.csv with snapshots in 2022, 2023, 2024 and 2025. The 2022 capture is in data/generic_ballot/generic_ballot_polls_2022.csv — 1,207 polls with pollster, field dates, sample size and D/R shares. The 2024 capture was not retrieved because archive.org went offline mid-session; the CDX query to repeat is recorded above.

Retrieved, 2026-09-25. The archive came back up. The CDX listing for the endpoint has snapshots through 2025-03-05, and the last useful one is 20241127022426 (42kB, taken after the 2024 election, so it covers that whole cycle). The 2025 captures are 1.6kB and 3.4kB — the endpoint was already returning almost nothing before 538 was shut down in March 2025.

scripts/ingest_generic_ballot.py fetches the captures through the provenance layer rather than reading a file out of data/, so every poll traces to a raw_artifact row holding the bytes that were retrieved. The polls land in fundamental_series as one row per poll — the generic ballot is a national quantity with no contest attached, so it cannot go in poll_question, which requires a race — with the field dates as the reference period and the publication date as the vintage. FundamentalsRepository.series_window reads them vintage-correct, because a poll published after the forecast date is a poll the forecast could not have seen.

CORRECTION, 2026-09-25: there were two files, and this section described one of them. 538's own polls/README.md says it plainly — "Current polls files contain data since the most recent election. Historical files contain data prior to the most recent election" — and names polls-page/data/generic_ballot_polls_historical.csv. That file, captured at 20250118200335, holds 4,832 national generic-ballot polls from 2016 to 2024, across the 2018, 2020, 2022 and 2024 cycles. Everything above was written after reading only the current file, and the conclusion that the corpus began in November 2020 was a statement about the lookup, not about the data (FM-63).

The corpus is now 3,815 polls over four cycles rather than 1,055 over one and a half, which is what makes an environment drift law estimable at all: experiments/environment_drift.json fits Var(h) = 0.112·h^0.59 from 371 non-overlapping increments of the generic-ballot path, and a law fitted on one cycle's path would have been a description of that cycle.

Two further corrections to what this section claimed.

The 2022 capture was truncated. 20221109051953, taken the day after the 2022 election, parsed to 388 polls whose newest was fielded to 3 August. So was 20221013105339. Every Wayback copy of the current file is cut off mid-body, which is why the September and October 2022 polls were missing and were put down to the source. ingest_generic_ballot.py now refuses a capture whose newest poll is more than 21 days from both the capture date and the election before it (FM-62). The local data/generic_ballot/generic_ballot_polls_2022.csv is a complete copy of that file and is not used: it has no provenance, and a CSV in data/ cannot say where it came from.

The rows are questions, not polls. One survey appears up to six times in these files — registered voters, likely voters, variant screens — so 4,832 rows are 3,815 polls. The ingest now takes one record per poll_id, preferring likely voters, then registered voters, then adults (FM-64).

Still missing, and named rather than smoothed over: generic-ballot polls fielded between 12 October and 8 November 2022. No capture of either file covers them. They are fillable from pollsters' own releases, which is what ingest_generic_ballot_direct.py does for the current cycle, and its --since flag is what would extend it backwards.

CORRECTION, 2026-09-25: the paragraph below was wrong, and wrong in the way this project has a failure mode for. I checked FiveThirtyEight's dead endpoint and a handful of guessed Wikipedia titles, found nothing, and reported a failed lookup as an absence — FM-43 and FM-46 again, a third time, and the owner caught it a second time. I never ran a Wikipedia search, and I never looked at a pollster's own site.

There are 680 individual generic-ballot polls for the 2026 cycle from 105 firms, spanning January 2025 to 22 September 2026: Emerson, Marquette, Quinnipiac, YouGov/Economist, Fox News, Echelon, AtlasIntel, McLaughlin, ActiVote and the rest. Every one is published by the firm that conducted it.

Routes, with the licence position on each.

route position
Decision Desk HQ prohibited. Terms of Use forbid "crawling, scraping, spidering or harvesting of any page, Content, personal information, or data". robots.txt allows /polls/, which does not override the Terms. Nothing was stored; the harvester written against it was deleted.
FiftyPlusOne 538-schema CSV API (generic.csv, house_general.csv, governor_general.csv, president_general.csv). Personal, non-commercial use with required attribution, behind a registered key. Registered as attribution, disabled: a standing restriction on what this project may do with its own outputs is a real cost, and the registration is the operator's act.
RealClearPolitics already prohibited.
The pollsters themselves the chosen route. No registration, no key, no terms acceptance, no restriction on downstream use. ADR-0009 already makes the pollster's own release authoritative and everything else a discovery layer to be checked against it, so reading the release directly collapses discovery and verification into one step.

src/elab/ingest/pollster_direct.py holds one extractor per firm, because each writes its release in its own prose. Seven are read — Quinnipiac, Rasmussen, The Economist/YouGov, Marquette, Emerson, Cygnal and Echelon Insights — each returning field dates, sample size, population and a cross-check against whatever else the release says about its own numbers. MISSING_EXTRACTORS names the six firms seen publishing a generic ballot and not yet read, with a reason each, because "no poll exists" and "nobody has written the twenty lines that read this firm's prose" are the distinction this project keeps getting wrong.

Echelon Insights is the safest read of the seven, and the reason is the document. They publish a topline: every question with its answer percentages, labelled by party in words, so there is no orientation to infer and no chart to read. The sub-items are the cross-check — "Definitely", "Probably" and "Lean" must reproduce each party's leaner-inclusive total, which is arithmetic the pollster did rather than a sentence they wrote. Their pages carry no date and their feed does, which is why a source may now declare a published_index: knowability is publication (FM-02) and this firm states it in one place only. Seven polls, April to September 2026.

Cygnal is the first source here whose numbers are not on a web page. Its poll page carries no vote shares at all; the public slide deck it links carries them, so reading it needed a second fetch and elab.parse.pdf_text. Three of the eight 2026 national releases state the ballot in prose with their own margin ("they lead 49% to 42%, a D+7 margin"), which is both the cross-check and the thing that decides which share is which; the other five describe movement without stating a level and produce nothing. The deck's chart page is not used although it is the obvious place to look — page 22 of the July deck extracts as eight percentages from two panels with nothing to say which is the topline, and a number read off that would be right some months. Provenance points at the deck, not the page: an archived page that does not state a figure is not evidence for that figure.

That also settled a reason that was being given for four other firms — and then checking the licences before writing the parsers changed two of them from technical to final (FM-77). "The release is a PDF" was the first obstacle recorded and not the operative one:

firm recorded reason what the licence check found
Fox News "the release is a PDF of toplines" prohibited. Terms of Use: "you may not copy, download, stream, scrape, capture, reproduce, duplicate, archive". robots.txt permits the pages and static.foxnews.com permits /foxnews.com/* where the PDFs live; neither overrides the Terms.
McLaughlin & Associates "monthly national poll is a PDF" prohibited. Terms restrict "any data mining, data harvesting, data extracting or any other similar activity". No robots.txt is served, which is not permission.
AtlasIntel "reports are PDFs and social posts" unverified. The terms page renders client-side and served no readable terms, so there is nothing to read a permission or a prohibition out of. Not fetched.
Echelon Insights "published as a deck" true of the deck, and their newsletter is silent too — but they also publish a topline, which is now read. The entry described the documents somebody had looked at rather than the ones the firm publishes.

All three are registered in config/sources.yaml with the quoted language. What remains a purely technical barrier is a form (Morning Consult), and two are refusals by the server: ActiVote answers a bot challenge and Quantus Insights serves a certificate that does not verify.

robots.txt checked 2026-09-25 for Quinnipiac (no Disallow rules at all), Emerson (one rule), Marquette (36, none blanket), Cygnal (/wp-admin/, /wp-includes/, /CLIENTS/, /dev/ — neither /polls/ nor /wp-content/uploads/), Echelon and Fox News (preview and search paths only). None of them restricts its published releases.

The superseded paragraph, kept because being wrong in public is part of the record:

2026 has no generic ballot, and this is now a settled conclusion rather than a pending search. The source stopped publishing before the cycle began. What remains on Wikipedia's 2026 House page is a table of aggregators — Decision Desk HQ, FiftyPlusOne, RealClearPolitics — which is not an acceptable input twice over: RealClearPolitics is disabled in the source registry, and an aggregate is another forecaster's model, which this project treats as a benchmark and never as ground truth. Also checked and absent: a "Nationwide opinion polling for the 2026 House elections" article, a Generic ballot article, individual-poll sections or labelled transclusions on the 2026 House page (only GenericBallotAgg exists), and the fivethirtyeight/data repository's congress-generic-ballot/ and polls/ folders, which the GitHub contents API shows hold one 1995–2016 trendline and presidential averages respectively.

Ingested, 2026-09-25: 1,055 individual polls spanning 2020-11-21 to 2024-11-04. Two captures with verified provenance, in fundamental_series with the field dates as the reference period, the publication date as the vintage, and the sample size in observation_n (migration 0022 — added after checking the spread rather than assuming it: 813 at the 5th percentile and 17,010 at the 95th, an interquartile ratio of 11.7, so equal weighting would throw away most of the information in the large polls).

Two things about the archive worth knowing before repeating this.

538 rotated the file per cycle. A capture from December 2023 holds 411 rows, all cycle 2024, starting the week after the 2022 election. There is no single file with every cycle in it.

A capture's timestamp is not the content's date. The snapshot at 20221109051953 — the day after the 2022 election — contains 454 rows ending 8 August 2022. The archive crawled a cached copy. So the complete 2022-cycle file, which must have carried September and October polls, exists only in captures taken between the 2022 election and the rotation, and the CDX endpoint returns 504 on every query narrowed to that window. Retry:

https://web.archive.org/cdx/search/cdx?url=projects.fivethirtyeight.com/polls/data/
generic_ballot_polls.csv&output=json&fl=timestamp,length&filter=statuscode:200
&from=20221110&to=20230601&collapse=digest

Consequence, stated rather than smoothed over. The 2022 cycle has generic-ballot coverage only through 8 August, three months before the election. Environment.staleness_days makes that visible and estimate returns UNAVAILABLE rather than a stale number at the 2022 election date, which is FM-45's lesson working rather than failing.

Not ingested: data/generic_ballot/generic_ballot_polls_2022.csv, 1,207 rows covering the full 2022 cycle, written by an earlier session with no raw_artifact row — its sha256 is not in the archive and which capture it came from is unknown. Bytes whose provenance is "somebody downloaded this" do not enter the corpus.

Priority note. This mattered less than expected. house_prior_too_wide.md shows the 2016 House forecast received a national swing of +4.63 against a realised +4.82 — essentially exact — and still missed by 38 seats. The environment is a real gap for 2018 onward and it is not what breaks the House forecast.

MEDSL per-state precinct files: what each state's file actually contains

Read for the 2024 cycle on 2026-09-25, while building district presidential baselines and certifying House results. These are not complaints about MEDSL, which publishes what the states publish; they are the shapes a consumer has to handle, and each was found because a number that looked reasonable disagreed with something.

state what the file does consequence here
OH lists every precinct under all fifteen districts no precinct-to-district crosswalk exists; taking the first district per precinct built one "district" holding the statewide vote, margin −11.3, entirely plausible (FM-70)
IN reports some votes both per precinct and inside a county-level batch totals come to 105% of the state's certified two-party vote; the margin is unaffected, which is why a margin check cannot see it
NC 4.2M of 5.7M votes are EARLY VOTING under county-wide pseudo-precincts dropping them lost 904k votes and tilted the state five points; they are apportioned across the county's districts instead
NY mixes a TOTAL mode row with the mode breakdown in 97 of 657 precincts summing both double-counted; the rule is TOTAL-if-present (FM-67)
GA, SC, ID same TOTAL-plus-modes mixture, statewide Georgia's total was exactly twice the truth, which is what that bug produces when every precinct does it
NE labels every presidential row NONPARTISAN party-based aggregation classified nothing; the nominees' surnames come from certified results as a fallback
WA presidential rows are partial 100% of precincts matched and the state still summed to D+5.02 against a certified D+18.93
NJ presidential rows are partial, districts' turnout ratios run 0.46× to 1.79× excluded from baselines and from House certification
LA about half the parishes present margin lands within a point of certified by luck, which is why vote coverage is checked as well as margin
OR labels a Constitution Party candidate DEMOCRAT (party_simplified) OR-02's two-party share disagrees by 2.1 points; this project's record is the one that matches the ballot
AZ, KS, MS assigned votes do not reproduce the certified margin within a point excluded from baselines, reason recorded

The checks that catch these live in scripts/build_district_baselines.py and are worth stating because the order they were added in is the order the failures were found: precinct match rate (caught OK, RI), statewide margin agreement (caught WA, NJ, AZ, KS, MS, LA), vote coverage floor and ceiling (caught LA's luck and IN's duplication), and crosswalk unambiguity (caught OH). Each one caught something the previous ones could not.