Skip to main content
Deliberate AcademyProfessional AI Education
AI Glossary

Benchmark

What is Benchmark?

A standardised test or dataset used to evaluate and compare AI model performance. Common benchmarks include MMLU (knowledge), HumanEval (coding), and HellaSwag (commonsense reasoning). Often cited in model announcements to justify capability claims.

Example in practice

A procurement manager evaluating two AI coding tools would compare their HumanEval scores to get an objective signal on code-generation quality before committing to a vendor.