An eval (short for evaluation) is a structured test that measures whether a language model or AI system produces the right output for a given input. Evals range from simple automated checks to human-scored assessments, and they are the primary way teams know whether a model is behaving as intended.
Why evals exist
Language models do not fail in the way deterministic software fails. A function that returns the wrong integer is clearly broken. A model that returns a plausible but subtly wrong answer is harder to catch. Evals provide a repeatable method for surfacing those failures before they reach production.
Without evals, the only signal is user complaints or manual inspection, neither of which scales.
What an eval contains
Most evals share the same structure:
- Input: a prompt, a document, a query, or a conversation turn
- Expected output: the correct answer, a reference response, or a rubric describing what good looks like
- Scoring method: exact match, semantic similarity, a secondary model acting as judge, or human review
The set of inputs used in an eval is called an eval dataset or benchmark. A well-maintained eval dataset grows over time, particularly when real failures are added back in after they are discovered.
Types of eval
Automated evals compare model output to a fixed reference. They are fast and cheap. They work well when the correct answer is unambiguous: factual retrieval, structured extraction, classification.
Model-graded evals use a second language model to score the output of the first. They handle open-ended tasks where exact matching is impractical. The tradeoff is that the grading model can itself be wrong, so its reliability needs to be established separately.
Human evals involve people rating or comparing outputs. They are the most reliable signal but the most expensive to run. Human evals are typically used to calibrate automated evals or to assess tasks where no automated proxy is good enough.
Evals in a RAG or agent system
In a RAG pipeline, evals often decompose into two separate questions: did the retrieval step return relevant context, and did the model use that context correctly? Conflating them makes it hard to diagnose failures.
In agentic systems, evals become more complex because the system takes a sequence of actions rather than producing a single response. The eval must account for intermediate steps, not just the final output.
What makes an eval useful
An eval is only as useful as its dataset and its scoring method. Common failure modes:
- Dataset leakage: the model has seen the eval inputs during training, making scores misleadingly high
- Proxy collapse: the metric being measured drifts from what actually matters
- Coverage gaps: the eval tests narrow conditions and misses the cases that fail in production
- No regression tracking: scores are checked once but not monitored across model versions or prompt changes
Teams that take evals seriously treat the eval suite as a long-lived engineering artefact, not a one-off check.
Evals and prompt changes
Every time a prompt changes, the model's behaviour can change. Evals serve as a regression suite for prompt engineering in the same way unit tests serve as a regression suite for code. This is one reason prompt engineering is less ad hoc than it is sometimes portrayed: the discipline only becomes manageable when there are evals to verify that a change improved things overall and did not silently break something else.
What we test for
For AI engineer candidates, evaluation design is one of two additional areas we assess beyond the standard vetting sessions. In the AI-native assessment, we look at whether a candidate can construct a valid eval for a realistic task, choose a scoring method that fits the output type, and explain where the eval might mislead. Closely related is knowing when a model is the wrong tool entirely. Full detail on how both sessions are structured and scored is at how we vet.
Short answers
What is the difference between an eval and a benchmark?
A benchmark is a published, standardised eval dataset used to compare models broadly. An eval is the general term, and includes private, task-specific test suites built by a team for their own system. Most production teams rely on internal evals rather than public benchmarks.
How many examples do you need in an eval dataset?
Estimates vary widely. Enough to cover the main input types and edge cases without being so small that a single outlier skews results. Hundreds of examples is a reasonable starting point for a focused task; broader systems need more. Quality and coverage matter more than raw count.
Can you use a language model to judge another language model's output?
Yes, and it is common practice for open-ended tasks. The grading model needs to be validated against human judgement first, and its own biases, such as preferring longer or more confident answers, should be understood before trusting its scores.