Phil Wagner

Staff Learning Designer and Technical Writer
Technical Lead of AI Enablement Education

Retrieval-augmented answers from a century-old magazine archive

Generated, then checked by the eval suite. Never verified against its source.

You want to see retrieval-augmented generation actually built, not described: a real corpus, a real index, and answers that cite the page they came from. Ask it something below, or read how it works underneath.

Ask the archive first if you would rather try it than read about it; how it works follows.

Sixteen issues of Popular Mechanics, 1905 to 1922, scanned by the Internet Archive and in the public domain. Ask a question and the system searches every page, then answers only from what it finds, with a link back to the scan. If the issues do not answer a question, it says so instead of guessing.

That is the whole claim of retrieval-augmented generation: a model's answer traced back to a real document, not to what it remembers. This page is a small, complete build of that claim, and the rest of it explains how.

Ask the archive

Pick an example for an instant answer, or type your own question.

Questions are logged with no name, address or account attached, so the archive's own team can see which questions it answers badly. Scans stay hosted at the Internet Archive; nothing here copies them.

How it works

A question either matches a cached example and is answered with no call, or it goes through a rate-limiting proxy to the archive: hybrid retrieval merges a keyword search and a meaning search into one ranked list, and a model answers only from the top of that list, with a citation, or says it does not know.

Two searches run on every question, not one. A keyword search finds the exact words a 1912 writer used. A meaning search finds a page that answers the question in different words. Each comes back as a ranked list, and the two lists are merged by rank rather than by score, because a keyword search's score and a meaning search's score are not measured the same way and cannot be compared directly.

The merged list goes to a model with one instruction: answer only from these pages, and say so plainly if they do not answer the question. That instruction is the difference between retrieval-augmented generation and a model guessing from what it was trained on. A wrong guess that sounds confident is worse than a stated "the issues don't say," so the second one is what this system is built to prefer.

The corpus

Sixteen issues, 2,356 pages, split into about 5,400 overlapping pieces of about 350 tokens each, roughly 260 words. A token is a word or a piece of one. A whole page is too long for most systems that turn text into a searchable meaning vector, so it is cut down first. The pieces overlap a little, so a sentence that starts near the end of one piece is not stranded there alone.

The meaning search here uses a light, local method rather than a hosted model, so this demo costs nothing to search and needs no network call until a question actually goes to the answer step. The same pipeline can swap in a heavier meaning search later without changing anything else, because chunking, search and answering are three separate steps.

Measured, not assumed

Retrieval changes here are measured against thirteen test questions. A script wrote each one from the archive's own text. Each was kept only after someone read the page it should find. For all thirteen, that page is in the top ten results. That score is called recall at ten.

One question still fails, and it is not one of the thirteen. Ask how kites are used, and pages that name kites once rank above the 1909 kite article. The system declines rather than answer from the wrong page. Fixing that ranking is open work, and it is why kites are not one of the examples above.

Keeping a small demo up and affordable

A hiring manager visits this page rarely, and a slow or broken page on that one visit costs more than it would on a busier site. Three small decisions carry that weight.

The example questions above cost nothing: their answers are checked once and stored on the page, not fetched again on every visit. Only a typed question reaches the archive, and it passes through a proxy that limits how often one visitor, or all visitors together, can ask before the archive's own daily budget is protected. The same layered limiting already runs elsewhere on this site's own pages, in front of a different model.

The archive's server also refuses to go live with a bad update. A new version is checked with a real question before it receives any traffic at all; if it fails, the version already running keeps serving and nothing is rolled back, because nothing was ever switched over.

What this demonstrates

A retrieval step, a grounding rule, and an honest refusal, built and measured rather than described. The same code is written to work on any collection of scanned documents, not only this one: a second, larger archive is already running on it, held back only by an open question about who may read it, not by the code.

Ask it without this page around it. The code and its measured results are not public yet; ask if you would like to see them.