Skip to content
Novistu

Glossary

Evals (evaluation suites)

Automated tests that measure AI output quality against your standards.

Evals are test sets and scoring pipelines that measure whether an AI system produces correct, useful outputs: accuracy on real cases, tone, format compliance, cost and latency. They run before releases and continuously in production, catching regressions from model updates or prompt changes. Without evals, AI quality is an opinion; with them, it is an engineering metric.

Example from practice

Your support agent's eval set of 200 real tickets must score above 95 percent grounded accuracy before any prompt or model change ships.

Related terms

Evaluating what these mean for your business? See the services built on them or ask us directly.