AI Benchmarks
AI benchmark rankings, model scores, and performance data.
Track live AI benchmark rankings, coding scores, math scores, and benchmark results across leading models from OpenAI, Anthropic, Google, Meta, DeepSeek, and more.
24 of 24 models
| # | Model | Org | Intelligence | Coding | Math | MMLU Pro | GPQA | LiveCodeBench | AIME 2025 | MATH 500 | SciCode | IFBench | HLE |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) | Anthropic | 53.4 | 81.6 | — | — | 9370.0 | — | — | — | 6310.0 | — | 5910.0 |
| 2 | Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Anthropic | 53.2 | 80.7 | — | — | 9340.0 | — | — | — | 6090.0 | — | 5870.0 |
| 3 | GPT-6 Astra (max) | OpenAI | 52.8 | 76.9 | — | — | 9610.0 | — | — | — | 5650.0 | — | 5470.0 |
| 4 | GPT-6 Astra (xhigh) | OpenAI | 52.5 | 75.9 | — | — | 9630.0 | — | — | — | 5570.0 | — | 5460.0 |
| 5 | Claude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback) | Anthropic | 51.2 | 79.1 | — | — | 9060.0 | — | — | — | 5870.0 | — | 5590.0 |
| 6 | GPT-6 Astra (high) | OpenAI | 51.0 | 77.1 | — | — | 9490.0 | — | — | — | 5540.0 | — | 5310.0 |
| 7 | Claude Opus 5 (Adaptive Reasoning, Max Effort) | Anthropic | 50.7 | 78.0 | — | — | 9320.0 | — | — | — | 5640.0 | — | 5490.0 |
| 8 | GPT-6 Astra (medium) | OpenAI | 49.7 | 76.7 | — | — | 9390.0 | — | — | — | 5420.0 | — | 5270.0 |
| 9 | Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) | Anthropic | 49.7 | 77.0 | — | — | 9370.0 | — | — | — | 5570.0 | — | 5440.0 |
| 10 | Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) | Anthropic | 49.7 | 76.5 | — | — | 9260.0 | — | — | — | 6100.0 | 6346.9 | 5550.0 |
| 11 | Claude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback) | Anthropic | 49.1 | 77.1 | — | — | 8860.0 | — | — | — | 5640.0 | — | 5380.0 |
| 12 | Muse Spark 1.3 (max) | Meta | 48.2 | 75.8 | — | — | 9350.0 | — | — | — | 5880.0 | — | 4870.0 |
| 13 | Claude Opus 5 (Adaptive Reasoning, High Effort) | Anthropic | 48.2 | 76.5 | — | — | 9370.0 | — | — | — | 5540.0 | — | 5280.0 |
| 14 | GPT-5.6 Sol (max) | OpenAI | 47.1 | 77.4 | — | — | 9410.0 | — | — | — | 5710.0 | 7265.3 | 4950.0 |
| 15 | Claude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback) | Anthropic | 47.0 | 75.2 | — | — | 8810.0 | — | — | — | 5670.0 | — | 4890.0 |
| 16 | GPT-6 Astra (low) | OpenAI | 46.0 | 75.7 | — | — | 9310.0 | — | — | — | 5410.0 | — | 4920.0 |
| 17 | Muse Spark 1.3 (xhigh) | Meta | 45.2 | 76.5 | — | — | 9410.0 | — | — | — | 5970.0 | — | 4750.0 |
| 18 | GPT-6 Astra (Non-reasoning) | OpenAI | 45.2 | 76.2 | — | — | 8950.0 | — | — | — | 5350.0 | — | 3710.0 |
| 19 | Claude Opus 5 (Adaptive Reasoning, Medium Effort) | Anthropic | 45.1 | 74.3 | — | — | 9190.0 | — | — | — | 5150.0 | — | 5130.0 |
| 20 | GLM-5.3 (max) | Z AI | 44.9 | 74.8 | — | — | 9170.0 | — | — | — | 5900.0 | — | 4230.0 |
| 21 | Grok 4.6 (high) | SpaceXAI | 44.4 | 76.8 | — | — | 9490.0 | — | — | — | 5650.0 | — | 4290.0 |
| 22 | Grok 4.6 (xhigh) | SpaceXAI | 44.3 | 75.9 | — | — | 9350.0 | — | — | — | 5300.0 | — | 4410.0 |
| 23 | GPT-5.6 Sol (xhigh) | OpenAI | 44.1 | 78.3 | — | — | 9310.0 | — | — | — | 5710.0 | 7102.0 | 4730.0 |
| 24 | Kimi K3 (max) | Kimi | 43.8 | 76.2 | — | — | 9350.0 | — | — | — | 5950.0 | — | 4690.0 |