This is the question running through every professional use of AI, and almost nobody answers it: how do you know it works? Usually the answer is "the answers look good". That is not a measurement, it is an impression, and it degrades without warning.
Definition
Evaluation means measuring the quality of an AI system's outputs on a set of known cases, reproducibly, so that two versions can be compared.
The important word is compare. An evaluation is not there to produce an absolute score, it answers one precise question: did this change improve or degrade the result?
Without evaluation, every modification becomes a bet. You change an instruction, the three answers you test look better, you ship. Three weeks later a case you never looked at has started failing, and nobody knows since when.
Building a case set, the only hard part
A useful evaluation fits in a table of twenty to fifty rows. Each row has a real input and what you expect as output.
Use real cases. Your actual customer requests, your actual documents. Invented examples produce an evaluation that validates a system which will fail in production.
Include what fails. This is the decisive point. A set made only of easy cases measures nothing. Ambiguous, out-of-scope or badly written requests are what separates two versions.
Write the success criterion. Not "good answer" but something checkable: contains the right amount, cites the right document, declines to answer, fits in three sentences.
Keep it frozen. A set modified at every test no longer allows comparison. It evolves, but deliberately, and you note when.
Three ways to judge, safest to most flexible
| Method | When to use it | Limit |
|---|---|---|
| Automatic check | The result is right or wrong | Only fits factual outputs |
| Human review | Writing quality, tone | Expensive, rarely replayed |
| Model as judge | High volume, written criteria | The judge has its own biases |
The third deserves caution: a model judging answers produced by a model tends to prefer a certain style. Use it to narrow down, not to settle an important choice.
When to replay the evaluation
At every instruction change. At every model change, including a version update at your provider, which can shift behaviour without you touching anything. And at regular intervals even without change, because your data evolves.
The classic mistake is evaluating on the same cases used to write the instruction. You then measure your ability to handle those specific cases, not the quality of the system. Hold back a set you never use to tune anything.
Frequently asked questions
Do you need a dedicated tool?
Not at first. A spreadsheet with your cases, a Workflow that replays them and a result column go a long way. Specialised tools earn their place when the number of cases exceeds what a spreadsheet keeps readable.
How many cases do you need?
Twenty well-chosen beat two hundred at random. The right measure is coverage: every type of request you receive should appear at least once, including the rare ones that cost you when they fail.
How do you evaluate a RAG system?
On two separate axes, otherwise you will not know what to fix. First retrieval: do the right extracts come back? Then writing: does the answer genuinely rest on those extracts? A RAG can fail at one without failing at the other.
Where do you start without losing a week?
By noting for a fortnight the cases where the system disappointed you: that notebook becomes your first evaluation set, and it is already representative. Our Claude Cowork course builds that practice in at go-live rather than after.