What's new

MMLU

knowledge

MMLU (Massive Multitask Language Understanding) tests a model across 57 academic subjects including STEM, humanities, social sciences, and more. It measures breadth of knowledge from elementary to professional level.

View paper / source

31

Models Tested

92.0

Best Score

87.0

Average Score

0–100

Scale Range

0.8x

Weight

How It Works

The model is given multiple-choice questions (4 options) across 57 subjects. Questions range from elementary mathematics to professional medicine and law. The test uses a few-shot format where the model sees examples before answering.

Why It Matters

MMLU is one of the most widely-cited benchmarks because it tests general knowledge breadth. A model that scores well on MMLU demonstrates broad competence across many domains, making it a useful proxy for general intelligence.

Limitations

MMLU relies on multiple-choice format which can be gamed. Some questions are ambiguous or have contested answers. It tests recall more than reasoning. Many modern models now saturate the benchmark (>90%), reducing its discriminative power.

Leaderboard — MMLU

# Model Provider Score
🥇 o3 OpenAI 92.0
🥈 o1 OpenAI 91.8
🥉 Grok 3 xAI 91.0
4 R1 DeepSeek 90.8
5 Gemini 2.5 Pro Preview 06-05 Google 90.5
6 GPT-4.1 OpenAI 90.2
7 GPT-4.5 OpenAI 90.2
8 DeepSeek V3 0324 DeepSeek 89.5
9 Claude Opus 4 Anthropic 89.0
10 GPT-4o (2024-05-13) OpenAI 88.7
11 Claude 3.5 Sonnet Anthropic 88.7
12 Llama 3.1 405B Meta 88.6
13 DeepSeek V3 DeepSeek 88.5
14 Claude Sonnet 4 Anthropic 88.0
15 Grok 2 xAI 87.5
16 GPT-4.1 Mini OpenAI 87.5
17 Mistral Large Mistral 87.0
18 Claude 3 Opus Anthropic 86.8
19 Gemini 2.5 Flash (batch) Google 86.5
20 Qwen2.5 72B Instruct Alibaba 86.1
21 Llama 3.3 70B Instruct Meta 86.0
22 Gemini 1.5 Pro Google 85.9
23 Llama 4 Maverick Meta 85.5
24 Nemotron 70B NVIDIA 85.0
25 Phi 4 Microsoft 84.8
26 Command A Cohere 83.5
27 Llama 4 Scout Meta 83.0
28 Claude 3.5 Haiku Anthropic 83.0
29 GPT-4o-mini (2024-07-18) OpenAI 82.0
30 Gemini 2.5 Flash Lite (batch) Google 80.0
31 Mistral Small 3.1 24B Mistral 78.0
All Benchmarks