Clef · Try the live demo →
Glossary

Evaluation (evals)

Evaluation is the practice of measuring an AI system against a fixed set of tasks so that every subsequent change is a measured delta rather than an impression. Without it a prompt change, a model upgrade or a retrieval tweak is indistinguishable from a mood. With it, the uncomfortable discoveries happen internally instead of in front of a user.

What a usable eval set looks like

Small and real. Twenty to fifty questions drawn from actual work, each with a known correct answer and a known source. Building it takes an afternoon with someone who does the job, and that afternoon is the most valuable part of most deployments — it forces the question of what "correct" means, which frequently turns out to be contested.

Larger sets look more rigorous and get maintained less. A set that runs in minutes and is trusted beats one that runs overnight and is quietly ignored after the second month.

Two kinds of measurement, and why you want both

Deterministic metrics cost nothing and catch regressions: does the answer contain the required anchors, how many distinct sources were used, is the structured output well-formed, how long did it take, what did it cost. These run on every change.

Judged metrics cost model calls and catch quality: coverage, depth, whether claims are supported, whether the answer is actually useful. These run less often, on a fixed rubric, with cached verdicts so that re-judging unchanged output is free. The discipline that matters is versioning the rubric — a score that moved because the rubric moved is worse than no score.

The failure the eval set is really guarding against

Silent regression on the long tail. A model upgrade improves the average and breaks a category — often the one with the specific formatting, the unusual document type, the edge case that someone depends on. Averages hide this by construction.

Which is why the set should over-represent the awkward cases rather than mirror the distribution of real traffic. It is not a sample; it is a tripwire.

What it is not to be confused with

Public benchmarks

Benchmarks compare models on shared tasks and say nothing about your corpus, your questions or your definition of correct. They are useful for narrowing a shortlist and useless for deciding whether a change to your system helped.

User feedback

Feedback arrives late, is biased towards the vocal, and cannot be replayed against a previous version. It is necessary and it is not a substitute: without an eval set you cannot tell whether the complaint reflects a regression or a preference.

Frequently asked

How large does the eval set need to be?+

Twenty to fifty tasks is enough to catch regressions and small enough to stay maintained. The composition matters far more than the size: include the cases that are awkward, the ones where the answer is genuinely not in the corpus, and the ones two experienced people would answer differently — that last group is diagnostic on its own.

Can a model judge its own output?+

For coarse distinctions, adequately; for fine ones, poorly, and with a bias towards fluent answers. The workable arrangement is deterministic metrics as the gate and judged metrics as the signal, with a human reviewing the judged cases that changed. A judge left entirely unsupervised drifts in the same direction as the system it judges.

When should evals be built?+

Before the first change to a working system, which in practice means during the pilot rather than after it. Built afterwards, the set encodes the behaviour you already have, including the parts that are wrong — it becomes a record of the status quo rather than a definition of correct.

Does this apply to a small deployment?+

Proportionately. A single-use-case assistant does not need a harness, but it does need twenty questions with known answers stored somewhere they can be re-run. That is the minimum that turns "it feels worse since last week" into a statement someone can check.

The afternoon that pays for itself

Sit with the person who does the work and write down twenty questions with their correct answers and sources. That document outlives every model you will use, and it is the only thing that makes "better" measurable.

Request a test