A course by Austin French

Evals in Production

A six-lesson course on testing LLM features after they ship: golden sets from real traffic, a judge you can trust, deciding what ships without a person, cost and latency, and evals in CI.

What it teaches

  1. Golden sets from production logs. Grow the golden set from real traffic. Scrub PII first, dedupe, save room for failures, label blind, and record where every case came from. TypeScript
  2. Calibrating the judge. Measure an LLM judge against human labels with kappa, read the disagreements, fix the rubric, and check for self-preference and position bias. TypeScript
  3. Autonomy gating. Decide which outputs ship without a person. Hard rules first, then a threshold picked from a measured error rate, a review queue, and a weekly audit. TypeScript
  4. The harness in Ruby on Rails. Move the eval loop into a Rails app so it tests the prompt you ship. A plain Ruby library, a rake task, a Minitest test and a table for history. Rails
  5. Cost and latency per feature. One record per model call, cost worked out from a price table when you read it, per-feature reports, budget alarms, and a model comparison on the golden set. TypeScript
  6. Shipping evals in CI. A cheap run on every PR that blocks only on cases the PR broke, a nightly full run that measures flakiness, and a PR comment people read. TypeScript

Who it's for

Engineers who have shipped an LLM feature, or are about to, and want to know whether a change made it better or worse. You should be comfortable reading TypeScript (five modules) and Ruby for the Rails module. It doesn't cover training models.

It picks up where the free LLM eval starter stops. The starter is a golden set, graders and a CI gate for a support-ticket triage feature. The course takes that same feature into production.

What's included

  • Six written lessons. Each one has the problem, the mental model, a worked example with real output, common mistakes, an exercise and a checklist.
  • Runnable code for every lesson, with tests. Five modules are TypeScript and one is a small Rails 8 app.
  • Synthetic data, labeled as synthetic, and a mock model, so every module runs offline on your machine without an API key. Modules 2, 4 and 5 can also call Anthropic or OpenAI with your own key. The Anthropic paths have each been run once, and the numbers are in the course's proof file; the OpenAI paths haven't.

Want it done for you?

The eval setup sprint on my hire page is this course applied to your feature: one week, a golden set from your real traffic, the harness in your repo, and a CI gate. The course is for doing it yourself.

Try it first

The starter harness is free and open source. Questions? Email me.