Definition
AI evaluation uses representative tasks and clear criteria to test outputs before and after release. It turns subjective impressions into evidence about quality and risk.
Why it matters
Without evaluation, teams cannot tell whether a change improved the system or quietly made it less trustworthy.
Business example
Before launch, a team scores 100 real support questions for citation quality, correctness, and escalation behavior.
When to use it
Use it before production and whenever prompts, models, data sources, or tools change.
When not to use it
Do not confuse a few impressive examples with a meaningful quality test.
How Automathing approaches it
We define the business-critical cases first, then set a practical quality bar and monitoring loop.
