← All work

Marginalia

Ask a question about one of twelve classic novels and get an answer that cites the text. Every quote is checked word for word against the passage it cites, and retrieval is measured, not guessed.

Marginalia is a reading companion for twelve novels everyone has heard of and fewer have finished: Pride and Prejudice, Moby-Dick, Dracula, Crime and Punishment, Jane Eyre, Gatsby and six more. Type a question. Get an answer with the passages it came from. Click a citation and the passage opens with the quote highlighted, and a link to read the whole chapter. Every quote has been checked against that passage before you see it, and the page tells you so.

It's also the clearest example I have of how I think RAG should be built. RAG (retrieval-augmented generation) means find the right passages first, then answer from them. Most demos stop at "it answered." This one measures the search, checks the quotes by code, shows its working on every answer, and tells you where it's weak.

A Marginalia answer: the sentence about live hedgehogs and flamingoes, a Verified quotes 3/3 badge, and the cited passage from chapter VIII open underneath with the quote highlighted
"What did the Queen use for croquet mallets and balls?" The answer, the badge, and the passage it came from with the quote highlighted. Made with no model at all. Ask it something yourself.

What it's for

  • Reading groups and students. "Which Bible story does Sonia read to Raskolnikov?" "What did Amy burn after quarreling with Jo?" Answers point at the page, so you can go read it.
  • Checking a half-remembered line. Ask in your own words. If the words don't match, it still finds the scene most of the time, and shows you which passages it looked at.
  • Seeing how RAG behaves. Every answer has a "how this answer was made" panel: the retrieved passages with their scores, how the answer was picked, and how each quote was checked. It's the panel I wish every RAG demo had.
  • Your documents. Swap the twelve novels for a product manual, a policy library or a contract set, and the same shape applies: measured retrieval, checked citations, a visible trace. That's the answers from your documents offer on my hire page.

Three steps

The model, when there is one, only writes. Retrieval and checking are code.
  1. Retrieve. Paragraph-based chunks of about 240 words that never cross a chapter. Search is BM25, a standard keyword ranking formula, plus a boost when query words sit next to each other, a boost for a book named in the question, and a share of the best neighboring chunk's score.
  2. Answer. With an Anthropic or OpenAI key, a model returns JSON where every claim carries exact quotes from numbered passages. Without a key, the answer is extractive: the sentences from the best passages that best cover the question.
  3. Check. The validator looks for each quote in the passage it cites, ignoring whitespace, quote-mark style, dashes and case. A quote that fails is shown as unverified, never dropped quietly. A model gets one retry with the failures listed.
The How this answer was made panel: eight retrieved passages with scores and BM25 ranks, then how the answer was picked and how quotes are checked
The "how this answer was made" panel. Every answer has one, and I open it more than the answer.

Measured, including where it's weak

The eval is 46 questions in four types (direct, paraphrased, no names, and names that appear in more than one book), each with its answer's passage found by searching the text. Then I added one change at a time and kept the changes that helped overall.

Found in top 1 Top 5 Top 10
Plain BM25, 160-word chunks 30.4% 52.2% 56.5%
Shipped config 34.8% 58.7% 65.2%
Shipped config, 12 held-out questions 41.7% 58.3% 58.3%

The improvement came from reading every miss, which is the part of this work I like most. The misses fell into two groups. In one, a scene names its people a paragraph before the event, so the chunk with the answer shares no words with the question. Bigger chunks and the neighbor score were aimed at that. In the other, the question is a pure paraphrase ("what advice did Nick's dad give him" against "criticizing anyone" and "my father"). Keyword search can't fix that without overfitting, and paraphrased questions still only find their passage in the top 10 a quarter of the time. That's what the optional vector search is for, and it's the first thing I'll measure with a key.

The held-out set is small: one question there is 8 points. These are numbers for comparing configs on this corpus, not an accuracy claim, and I'd rather say that than round up.

With a real model writing the answers (Claude Haiku 4.5, one run over all 58 questions), 62 of 67 quotes passed the check on the first reply. Three answers needed the retry, and two were shown with a quote flagged as unverified. The model said "not found" 19 times, and in 17 of those the right passage wasn't among the eight it was given, so it declined instead of inventing. When the right passage was there, a verified quote came from it 34 times out of 35. The whole run cost 25 cents. The code and the full table are in the repo.

The quote validator is tested the other way around. For each question, the eval plants fake quotes in the passages actually retrieved: invented sentences, one word changed, a real quote cited to the wrong passage, real pieces in the wrong order. It caught 320 of 320, and passed 136 of 136 real quotes. CI fails if either drops. What it can't catch is a real quote attached to a claim it doesn't support. It proves the words are in the book, not that the claim follows from them. 61 tests, no network.

Marginalia and Gutenberg MCP

Both check quotes against the text, and I built both. Marginalia is the app: it retrieves, answers and checks, for twelve books it has already prepared. Gutenberg MCP is a tool for someone else's assistant: it checks any quote against any of 75,000 books, and never answers anything itself. If you have a question, use this. If you want Claude to stop misquoting, use that.

What's next

  • Run it more than once, and on a second model. One run gives one number; the variance between runs is the next thing to know.
  • Turn on hybrid search and rerun the eval. The vector path is built; the numbers aren't.
  • More books, chosen by what people ask. The corpus build is one command.
  • Your documents instead of novels. The eval set comes from your real questions, and the validator checks quotes against your sources.