Evaluation: knowing whether your AI actually works

Evaluation means measuring the quality of an AI system's answers on a set of known cases, rather than by impression.
3 min read
Believemy logo

This is the question running through every professional use of AI, and almost nobody answers it: how do you know it works? Usually the answer is "the answers look good". That is not a measurement, it is an impression, and it degrades without warning.


Definition

Evaluation means measuring the quality of an AI system's outputs on a set of known cases, reproducibly, so that two versions can be compared.

The important word is compare. An evaluation is not there to produce an absolute score, it answers one precise question: did this change improve or degrade the result?

Good to know

Without evaluation, every modification becomes a bet. You change an instruction, the three answers you test look better, you ship. Three weeks later a case you never looked at has started failing, and nobody knows since when.


Building a case set, the only hard part

A useful evaluation fits in a table of twenty to fifty rows. Each row has a real input and what you expect as output.

Use real cases. Your actual customer requests, your actual documents. Invented examples produce an evaluation that validates a system which will fail in production.

Include what fails. This is the decisive point. A set made only of easy cases measures nothing. Ambiguous, out-of-scope or badly written requests are what separates two versions.

Write the success criterion. Not "good answer" but something checkable: contains the right amount, cites the right document, declines to answer, fits in three sentences.

Keep it frozen. A set modified at every test no longer allows comparison. It evolves, but deliberately, and you note when.


Three ways to judge, safest to most flexible

MethodWhen to use itLimit
Automatic checkThe result is right or wrongOnly fits factual outputs
Human reviewWriting quality, toneExpensive, rarely replayed
Model as judgeHigh volume, written criteriaThe judge has its own biases

The third deserves caution: a model judging answers produced by a model tends to prefer a certain style. Use it to narrow down, not to settle an important choice.


When to replay the evaluation

At every instruction change. At every model change, including a version update at your provider, which can shift behaviour without you touching anything. And at regular intervals even without change, because your data evolves.

Warning

The classic mistake is evaluating on the same cases used to write the instruction. You then measure your ability to handle those specific cases, not the quality of the system. Hold back a set you never use to tune anything.


Frequently asked questions

Question

Do you need a dedicated tool?

Not at first. A spreadsheet with your cases, a Workflow that replays them and a result column go a long way. Specialised tools earn their place when the number of cases exceeds what a spreadsheet keeps readable.


Question

How many cases do you need?

Twenty well-chosen beat two hundred at random. The right measure is coverage: every type of request you receive should appear at least once, including the rare ones that cost you when they fail.


Question

How do you evaluate a RAG system?

On two separate axes, otherwise you will not know what to fix. First retrieval: do the right extracts come back? Then writing: does the answer genuinely rest on those extracts? A RAG can fail at one without failing at the other.


Question

Where do you start without losing a week?

By noting for a fortnight the cases where the system disappointed you: that notebook becomes your first evaluation set, and it is already representative. Our Claude Cowork course builds that practice in at go-live rather than after.

Related terms

Discover our aI and automation glossary

The vocabulary of artificial intelligence and automation, explained for people who want to use it in their business, not for people who build the models.

Share this article

Want to help us? Share this article on your networks or even better: on your site, in an article or in your newsletter.