AI Glossary
Evaluation (AI)
What is Evaluation (AI)?
The systematic process of measuring an AI model's performance, accuracy, fairness, and behavior. Evaluation methods include automated benchmarks, human preference ratings, red-teaming, and domain-specific accuracy tests. Rigorous evaluation is essential before deploying AI in professional or high-stakes contexts.
Example in practice
A procurement team assessing an AI contract analysis tool would run it on 50 historical contracts with known issues, then measure its detection rate — this task-specific evaluation is more meaningful than published benchmark scores.
Learn more
See Evaluation (AI) applied in a professional context through this free course.
AI Strategy for Leaders →