AI Concepts
Benchmark
Definition
A standardised test used to evaluate and compare AI model performance. Common benchmarks include MMLU (general knowledge), GPQA (graduate-level science), HumanEval (coding), and MATH (mathematical reasoning). Benchmarks enable objective comparison between models.
In Plain English
A standardised exam for AI models. Just like students take SATs to compare performance, AI models take benchmarks. A model scoring 90% on MMLU means it answered 90% of a massive multiple-choice knowledge test correctly.