AI evaluation
How to evaluate an AI product
The short answer
Evaluate the product against the real task, not the model in isolation. Define what good looks like, build representative and adversarial cases, measure the failures that matter, set thresholds, and keep the evaluation suite as a regression test.
Start with the decision the product must support
Generic model benchmarks rarely answer whether a product is safe or useful in its operating context. Evaluation should begin with the target user task, the acceptable outcome, the consequential failure modes and what happens after the model responds.
Build a representative evaluation set
Include common cases, difficult cases, ambiguous inputs, missing information, edge cases and examples that previously caused failure. The set should represent the distribution the product actually sees rather than the easiest examples to score.
Measure more than answer quality
- Task success and failure severity.
- Unsupported or fabricated claims.
- Tool selection and action accuracy.
- Calibration and handling of uncertainty.
- Latency, cost and retry behaviour.
- Human review burden and override patterns.
Choose thresholds before launch
A metric is only useful if the team knows what level is acceptable. Thresholds should reflect consequence: a cosmetic wording failure and an unsafe action should not carry the same weight.
Keep evaluation in the delivery loop
Model changes, prompt changes, retrieval changes and tool changes can all create regressions. The evaluation suite should run repeatedly as the product evolves, with new failure cases added as they are discovered.