AI Benchmarks
AI benchmark rankings, model scores, and performance data.
Track live AI benchmark rankings, coding scores, math scores, and benchmark results across leading models from OpenAI, Anthropic, Google, Meta, DeepSeek, and more.
24 of 24 models
| # | Model | Org | Intelligence | Coding | Math | MMLU Pro | GPQA | LiveCodeBench | AIME 2025 | MATH 500 | SciCode | IFBench | HLE |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 (Adaptive Reasoning, Max Effort) | Anthropic | 63.1 | 78.0 | — | — | 9320.0 | — | — | — | 5570.0 | — | 5490.0 |
| 2 | Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) | Anthropic | 62.5 | 77.0 | — | — | 9370.0 | — | — | — | 5500.0 | — | 5440.0 |
| 3 | Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) | Anthropic | 62.1 | 76.5 | — | — | 9260.0 | — | — | — | 6020.0 | 6346.9 | 5550.0 |
| 4 | Claude Opus 5 (Adaptive Reasoning, High Effort) | Anthropic | 61.5 | 76.5 | — | — | 9370.0 | — | — | — | 5430.0 | — | 5280.0 |
| 5 | Grok 4.6 (high) | SpaceXAI | 60.9 | 76.8 | — | — | 9490.0 | — | — | — | 5360.0 | — | 4290.0 |
| 6 | GPT-5.6 Sol (max) | OpenAI | 60.9 | 77.4 | — | — | 9410.0 | — | — | — | 5610.0 | 7265.3 | 4950.0 |
| 7 | Kimi K3 (max) | Kimi | 59.7 | 76.2 | — | — | 9350.0 | — | — | — | 5870.0 | — | 4690.0 |
| 8 | GPT-5.6 Sol (xhigh) | OpenAI | 59.0 | 78.3 | — | — | 9310.0 | — | — | — | 5600.0 | 7102.0 | 4730.0 |
| 9 | Claude Opus 5 (Adaptive Reasoning, Medium Effort) | Anthropic | 58.6 | 74.3 | — | — | 9190.0 | — | — | — | 5070.0 | — | 5130.0 |
| 10 | Qwen3.8 Max | Alibaba | 58.1 | 71.8 | — | — | 9270.0 | — | — | — | 5290.0 | — | 4300.0 |
| 11 | Qwen3.8 2.4T A95B | Alibaba | 57.7 | 71.9 | — | — | 9350.0 | — | — | — | 5160.0 | — | 4240.0 |
| 12 | GPT-5.6 Sol (high) | OpenAI | 57.3 | 77.2 | — | — | 9280.0 | — | — | — | 5690.0 | 6918.4 | 4600.0 |
| 13 | Claude Opus 4.8 (Adaptive Reasoning, Max Effort) | Anthropic | 57.3 | 74.3 | — | — | 9200.0 | — | — | — | 5350.0 | 6224.5 | 4870.0 |
| 14 | Muse Spark 1.2 (xhigh) | Meta | 56.8 | 72.2 | — | — | 9040.0 | — | — | — | 5640.0 | — | 4550.0 |
| 15 | GPT-5.6 Terra (max) | OpenAI | 56.6 | 76.7 | — | — | 9250.0 | — | — | — | 5390.0 | 7122.4 | 4290.0 |
| 16 | GPT-5.5 (xhigh) | OpenAI | 56.3 | 74.9 | — | — | 9350.0 | — | — | — | 5610.0 | 7585.0 | 4580.0 |
| 17 | Gemini 3.7 Flash (high) | 56.0 | 76.1 | — | — | 9450.0 | — | — | — | 5680.0 | — | 4790.0 | |
| 18 | Grok 4.5 (high) | SpaceXAI | 55.8 | 72.4 | — | — | 9310.0 | — | — | — | 5410.0 | — | 4270.0 |
| 19 | GPT-5.6 Sol (medium) | OpenAI | 55.6 | 76.3 | — | — | 9260.0 | — | — | — | 5650.0 | 6959.2 | 4220.0 |
| 20 | Claude Sonnet 5 (Adaptive Reasoning, Max Effort) | Anthropic | 55.3 | 71.5 | — | — | 9110.0 | — | — | — | 5360.0 | — | 4130.0 |
| 21 | Claude Opus 4.7 (Adaptive Reasoning, Max Effort) | Anthropic | 55.0 | 73.6 | — | — | 9140.0 | — | — | — | 5450.0 | 5863.9 | 4230.0 |
| 22 | GPT-5.5 (high) | OpenAI | 54.7 | 71.6 | — | — | 9320.0 | — | — | — | 5590.0 | 7163.3 | 4500.0 |
| 23 | Gemini 3.7 Flash (medium) | 53.4 | 71.5 | — | — | 9210.0 | — | — | — | 5790.0 | — | 3900.0 | |
| 24 | Muse Spark 1.1 (xhigh) | Meta | 53.2 | 71.3 | — | — | 8980.0 | — | — | — | 5820.0 | — | 4620.0 |