Second Draft
A writing coach that points at exact sentences. Paste a draft, pick a goal, and get the three fixes that would change it most, each quoting a sentence you wrote.
Most writing feedback is vague. "Tighten this." "Show, don't tell." I wanted a coach that has to point. Paste a draft of up to 3,000 words, pick a goal (tighter, more vivid, clearer, more like you), and you get the three fixes that would change the piece most. Each one quotes a sentence you wrote, shows a before and after, says why it matters, and gives you a short exercise. Revise, run it again, and watch the numbers move.
Three fixes at a time is the whole idea. A wall of red makes people close the tab. Three things you can do today gets a second draft written.
What it's for
- Writers revising a draft. A story opening, an essay, a cover letter. Pick the goal and get three concrete moves.
- Teachers, who want feedback students can check against their own text instead of taking on faith.
- Anyone who writes for work. The "clearer" goal leans on passive voice, hedges, filler and dense sentences, which is most of what makes a memo hard to read.
- Any agent that edits text someone else wrote. The validator here is the reusable part: a quote has to be in the place it says it is, or it doesn't ship. That's the agent build on my hire page.
Deterministic below, probabilistic above
- Eleven detectors find facts with character offsets: passive voice, filler and hedges, adverbs, cliches, repeated words, dialogue tags, named emotions, sentence rhythm, paragraph shape, the opening line and dense sentences.
- Without a key, a deterministic coach ranks findings by a leverage score and every card shows its own arithmetic. Rewrites are built by code.
- With a key, an agent coaches. Claude or an OpenAI model runs a tool-use loop with five tools (analyze the draft, read findings, read sentences, compare drafts, submit), with a turn limit, a token budget and a timeout. Any failure falls back to the deterministic coach and says so on the page.
- Every citation is checked. An item only survives if its sentence number is real and its quote appears verbatim in that sentence. Offsets are computed from the draft, never taken from the model. Bad items get one repair turn, then they're dropped.
The eval, with the error list
The eval set is 25 synthetic passages (fiction, cover letters, essays, business notes, four clean controls) with 106 labeled problems, labeled from an editor's point of view before the detectors were run. The detectors find 86% of the labels, and 88% of what they flag is right.
That precision used to be 82%, and the way it got better is my favourite part of the project. Reading all 20 false positives, 12 came from the repeated-word detector: names at the start of a sentence, repetition a writer did on purpose ("every summer ... every summer"), and nouns that name the subject. Fixing the first two, dropping an opening-line rule whose 3 findings were all wrong, and one fix to passive voice got it to 12 false positives. Repeated words is still the weakest detector at 43% precision. Telling a clumsy echo from a needed noun takes meaning, which is the model's job, so it carries the lowest weight.
The citation validator was tested by planting 887 bad citations of the kinds a model plausibly makes. It catches all of them by construction, so the useful number was the other direction: on its first run it rejected 2 of 269 good citations, because a punctuation-only rewrite looked unchanged. That's fixed, with a test. 99 tests in all, including the agent loop against mocked Anthropic and OpenAI responses, turn by turn, with the 401, 429 and 5xx paths.
Then the agent ran for real: all 25 passages, once each, on three models. Claude Haiku 4.5 finished every passage without falling back, and 21 of its 25 first submissions passed the validator in full; the repair turn fixed the other four. Haiku 5.5 went 23 of 25 and never submitted a bad quote (its two misses were empty submissions), for three cents. Sonnet 5.5 went 22 of 25; on two fiction passages its first submission was missing required fields, and when I reran those passages on their own it passed all three first time. The validator is the reason none of that reached the page. The per-passage rows are in the repo.
The set is small, and I wrote both the passages and the detectors. It's a regression harness and an honest error list, which is what I'd want from a tool before I trusted it with my own drafts.
What's next
- Run it more than once. One run per model gives one number each; the variance between runs is the next thing to know.
- A house style as a goal. Feed it a few pieces you like and let "more like you" learn from them, with the same checks.
- Longer work, chapter by chapter, with progress across the whole thing instead of one draft.
- The same loop for other documents. A support reply, a report, a policy. Any agent that edits text someone else wrote should quote the line it's changing.