Hire me

I build the AI features that have to be right. Tools your assistant can call. Answers that cite the page they came from. Agents that show their work. Fixed scope, agreed in writing, and tested before you see it.

Custom MCP server

Your data, usable from Claude by the end of the week.

MCP (Model Context Protocol) is how Claude, Cursor and ChatGPT call outside tools. I build the server on the other end of it for your API, database or document store: a few tools shaped around the questions your people actually ask, a general one for everything else, and the validation, rate limits and plain-language errors that make it safe to leave running. You connect it once. After that, your team just asks.

  • 3 days Basic. Up to 3 read-only tools over your API, running locally, with install docs for Claude Desktop, Claude Code and Cursor.
  • 5 days Standard. Up to 8 tools, including writes that ask for confirmation first. Validation, tests and the playground.
  • 7 days Premium. A remote server on your cloud with auth (API key or OAuth), rate limits and a daily cap, plus docs and a handoff call.
time
3 to 7 days
price
Fixed scope and a fixed price, agreed in writing before any work starts.

You get

  • Tools designed from the real questions, not from your schema
  • Every input validated and escaped before it touches your system
  • Errors the model can act on, so it recovers instead of guessing
  • A web playground where anyone can run every tool and see the raw result
  • Install docs for Claude Desktop, Claude Code and Cursor

See it working

Answers from your documents

Ask a question, get an answer that cites the page it came from.

Search and answers over your own documents (RAG, retrieval-augmented generation), built so you can trust what comes back. Every quote in an answer is checked by code against the source it cites before anyone sees it. Retrieval is measured on real questions from your domain, so you know what it finds and what it misses before your users do. And every answer shows its working: which passages it read, how they scored, which quotes passed.

time
Scoped on a call
price
Fixed scope and a fixed price, agreed in writing before any work starts.

You get

  • Retrieval tuned on a question set from your domain, with the numbers
  • A quote check that catches invented, altered and misattributed quotes
  • A "how this answer was made" view on every answer
  • Works with Anthropic or OpenAI, or with no model at all for the cheap path

See it working

An agent that does the job

Tools on your systems, limits on everything, a record of what it did.

An agent that uses tools on your systems to get a piece of work done, end to end. It runs with limits on turns, tokens and time, falls back to something safe when the model fails, and anything it cites is checked by code before a person sees it. The model decides what matters and explains it. Code does the measuring, and code holds the line.

time
Scoped on a call
price
Fixed scope and a fixed price, agreed in writing before any work starts.

You get

  • The tools, scoped to what the job needs and nothing more
  • Turn, token and time budgets, with a safe fallback that says it fell back
  • A validator between the model and the person
  • Tests that script the model's side of the conversation, including the failures

See it working

Eval setup sprint

One week to get your LLM feature under test, in your repo.

You've shipped an LLM feature and you're changing the prompt by feel. In one week I build the test suite for it: a golden set from your real traffic, graders that score the same way every time, a diff that shows which cases changed between two versions, and a gate in CI that blocks the merge when quality drops. TypeScript or Rails.

time
One week
price
Fixed scope and a fixed price, agreed in writing before any work starts.

You get

  • A golden set with where each case came from, scrubbed of personal data
  • Code graders for everything that has a right answer, a model judge only for what doesn't
  • Per-tag floors, so a strong average can't hide a category that collapsed
  • The CI gate, and a short write-up of what the first run found

See it working

Or learn it: Evals in Production

The eval sprint as a course. Launching soon.

A six-lesson course on testing LLM features after they ship: golden sets from real traffic, a judge you can trust, deciding what ships without a person, cost and latency, and evals in CI. 6 written lessons, each with code you run, in TypeScript and Rails. It starts from the free LLM eval starter and takes the same feature into production.

See the course

Why it holds up

One rule runs through every project: code decides what's true, and the model works on top of it.

The bottom row is code with tests. The top row is where the model earns its keep. The validator is what lets a model's sentence reach a person.

Code decides what's true

What the data says, what the engine says, whether a quote is in the text: code answers those, and it's tested. The model ranks, explains and writes on top. That's why these tools don't make things up.

Every claim points at its source

If an answer quotes something, code has confirmed the quote is there. If a number is in a write-up, it came from a command you can run again.

You can watch it work

Every project ships with a playground that runs the real tools and shows the raw result next to the friendly one. Your team can see exactly what the model had to work with.

Fast to build, slow to read

AI tools, mostly Claude Code, are part of how I build, which is why a week is enough. Everything they write is read, tested and measured before it ships.

Next step

Tell me what you're building and what's in the way. Book a 30-minute call or email me.