usnewsmap.com searches 23 million pages of historical American newspapers and maps every match by where and when it was printed. Play the timeline to watch a word, a name or a story spread across the country.
The pages come from Chronicling America, the Library of Congress and National Endowment for the Humanities collection of digitized newspapers, 1736–1963. This is a rebuild of the original US News Map (2015–2016, Georgia Tech Research Institute and eHistory.org at the University of Georgia), whose code is preserved at tgoodyear/usnewsmap.
- Phrase, all-words, any-word and proximity search over the OCR text, with filters by date, state and language.
- Complete counts for every matching page, by place and time, in one cacheable response; playback and cumulative views run in the browser.
- A relative-rate view: how much more or less a place printed a term than its digitized pages predict.
- Snippets with a link to each page at the Library of Congress, a table view, keyboard playback, and a shareable URL for every search.
- A public JSON API under
/v1, the same one the site uses (design doc 06).
A Rust API (axum) and a Quickwit searcher run side by side in one Azure Container App, and the API also serves the site. A Rust ingest pipeline, run as Container Apps jobs, turns the Library of Congress's bulk OCR into curated Parquet on Blob Storage, builds the search indexes from it and publishes versioned snapshots. Everything is defined in Bicep, every service signs in with an Entra ID managed identity rather than a key, and there are no servers to maintain. The architecture overview has the details.
Needs a stable Rust toolchain (rust-toolchain.toml) and Node.js 24. No Azure account or newspaper data: the API serves a small synthetic corpus from fixtures/.
cargo run -p usnm-api # terminal 1: the API on http://localhost:8080
cd web && npm ci && npm run dev # terminal 2: the site on http://localhost:5173cargo test --workspace runs the unit and API tests against the same fixtures. The development guide covers running against Quickwit, the ingest pipeline and the local Azure stand-ins.
| Path | What |
|---|---|
crates/ |
The Rust workspace: usnm-core (domain types and the query language), usnm-search (search backends), usnm-store (Blob Storage), usnm-api (the API, which also serves the site) and usnm-ingest (the pipeline) |
web/ |
The web app: React, MapLibre GL and deck.gl (README) |
infra/ |
The Azure infrastructure in Bicep, one deployment stack per environment, and the Quickwit configs (README) |
ja-ocr/ |
The Japanese OCR job (Python, NDLOCR-Lite) |
scripts/ |
Stand an environment up, deploy and tear it down; read its logs; local Azure stand-ins; load tests |
fixtures/ |
A small synthetic corpus for development and tests (README) |
ops/ |
Saved log queries and the history of published index versions |
docs/ |
Guides, the design documents and the technical notes (index) |
- Development: running and testing locally.
- Configuration: the API's settings.
- Operations: running the ingest pipeline in Azure.
- Infrastructure: what the Bicep deploys and how to stand an environment up.
- Design documents and architecture decision records: what was built and why.
- Technical notes: dated write-ups of work on the index, the OCR and the corpus's coverage, with the measurements behind each change.
See CONTRIBUTING.md. To report a vulnerability, follow SECURITY.md and don't open a public issue.
- Newspaper pages, OCR text and title records, the source of most place coordinates: Chronicling America, from the National Digital Newspaper Program of the National Endowment for the Humanities and the Library of Congress.
- The original US News Map (2016): eHistory.org at the University of Georgia (Claudio Saunt and Steve Berry) and the Georgia Tech Research Institute (Trevor Goodyear, David Ediger and Zach Suffern).
- Basemap: OpenFreeMap, with map data from OpenStreetMap contributors.
- Built on Quickwit, MapLibre GL JS and deck.gl.
The code is released under the MIT License. The newspaper content belongs to its sources: see the Library of Congress's rights and access statement for Chronicling America.
