← All ridiculous benchmarks

Version Bench

96 of 96
Version Bench96 dated model releases. Advertised version.Version BenchSix is higher than five. The methodology is exhaustive.Advertised version1.02.53.64.14.75.15.56.1Nov 2022Nov 2023Nov 2024Oct 2025Oct 2026GPT-3.5 Turbo · 3.5 · Nov 30, 2022GPT-4 · 4 · Mar 14, 2023GPT-3.5 Turbo (0613) · 3.5 · Jun 13, 2023Claude 2.0 · 2.0 · Jul 11, 2023GPT-4 Turbo · 4 · Nov 6, 2023Claude 2.1 · 2.1 · Nov 21, 2023Gemini 1.0 Pro · 1.0 · Dec 6, 2023Claude 3 Haiku · 3 · Mar 4, 2024Grok-1 · 1 · Mar 17, 2024GPT-4o (May '24) · 4 · May 13, 2024Gemini 1.5 Flash (May '24) · 1.5 · May 14, 2024Gemini 1.5 Pro (May '24) · 1.5 · May 15, 2024Claude 3.5 Sonnet (June '24) · 3.5 · Jun 21, 2024GPT-4o mini · 4 · Jul 18, 2024GPT-4o (Aug '24) · 4 · Aug 6, 2024o1-mini · 1 · Sep 12, 2024Gemini 1.5 Flash (Sep '24) · 1.5 · Sep 24, 2024Gemini 1.5 Flash-8B · 1.5 · Oct 3, 2024Claude 3.5 Haiku · 3.5 · Oct 22, 2024GPT-4o (Nov '24) · 4 · Nov 20, 2024o1 · 1 · Dec 5, 2024Gemini 2.0 Flash (experimental) · 2.0 · Dec 11, 2024Grok 2 (Dec '24) · 2 · Dec 12, 2024GPT-4o mini Realtime (Dec '24) · 4 · Dec 17, 2024Gemini 2.0 Flash Thinking Experimental (Dec '24) · 2.0 · Dec 19, 2024Gemini 2.0 Flash Thinking Experimental (Jan '25) · 2.0 · Jan 21, 2025o3-mini · 3 · Jan 31, 2025Gemini 2.0 Flash (Feb '25) · 2.0 · Feb 5, 2025GPT-4o (ChatGPT) · 4 · Feb 15, 2025Grok 3 · 3 · Feb 19, 2025Claude 3.7 Sonnet (Non-reasoning) · 3.7 · Feb 24, 2025Gemini 2.0 Flash-Lite (Feb '25) · 2.0 · Feb 25, 2025GPT-4.5 (Preview) · 4.5 · Feb 27, 2025o1-pro · 1 · Mar 19, 2025Gemini 2.5 Pro Preview (Mar' 25) · 2.5 · Mar 25, 2025GPT-4o (March 2025, chatgpt-4o-latest) · 4 · Mar 27, 2025GPT-4.1 · 4.1 · Apr 14, 2025o3 · 3 · Apr 16, 2025o4-mini (High) · 4 · Apr 16, 2025Gemini 2.5 Flash Preview (Non-reasoning) · 2.5 · Apr 17, 2025Gemini 2.5 Pro Preview (May' 25) · 2.5 · May 6, 2025Gemini 2.5 Flash (Non-reasoning) · 2.5 · May 20, 2025Claude 4 Opus (Non-reasoning) · 4 · May 22, 2025Gemini 2.5 Pro · 2.5 · Jun 5, 2025o3-pro · 3 · Jun 10, 2025Gemini 2.5 Flash-Lite (Non-reasoning) · 2.5 · Jun 17, 2025Grok 4 · 4 · Jul 10, 2025Claude 4.1 Opus (Non-reasoning) · 4.1 · Aug 5, 2025GPT-5 (High) · 5 · Aug 7, 2025Grok 4 Fast (Non-reasoning) · 4 · Sep 19, 2025GPT-5 Codex (High) · 5 · Sep 23, 2025Gemini 2.5 Flash-Lite Preview (Sep '25) (Non-reasoning) · 2.5 · Sep 25, 2025Claude 4.5 Sonnet (Non-reasoning) · 4.5 · Sep 29, 2025Claude 4.5 Haiku (Non-reasoning) · 4.5 · Oct 15, 2025GPT-5.1 (High) · 5.1 · Nov 13, 2025Gemini 3 Pro Preview (High) · 3 · Nov 18, 2025Grok 4.1 Fast (Non-reasoning) · 4.1 · Nov 19, 2025Claude Opus 4.5 (Non-reasoning) · 4.5 · Nov 24, 2025GPT-5.2 (Xhigh) · 5.2 · Dec 11, 2025Gemini 3 Flash Preview (Non-reasoning) · 3 · Dec 17, 2025Claude Opus 4.6 (Non-reasoning, High) · 4.6 · Feb 5, 2026Gemini 3 Deep Think · 3 · Feb 5, 2026GPT-5.3 Codex (Xhigh) · 5.3 · Feb 5, 2026Claude Sonnet 4.6 (Non-reasoning, High) · 4.6 · Feb 17, 2026Gemini 3.1 Pro Preview · 3.1 · Feb 19, 2026Gemini 3.1 Flash-Lite · 3.1 · Mar 3, 2026GPT-5.4 (Xhigh) · 5.4 · Mar 5, 2026Grok 4.20 0309 (Reasoning) · 4.20 · Mar 10, 2026GPT-5.4 mini (Xhigh) · 5.4 · Mar 17, 2026Grok 4.20 0309 v2 (Reasoning) · 4.20 · Apr 7, 2026Claude Opus 4.7 (Max) · 4.7 · Apr 16, 2026GPT-5.5 (Xhigh) · 5.5 · Apr 23, 2026Grok 4.3 (High) · 4.3 · Apr 30, 2026GPT-5.5 Instant (May 2026) · 5.5 · May 5, 2026Gemini 3.5 Flash (High) · 3.5 · May 19, 2026Claude Opus 4.8 (Max) · 4.8 · May 28, 2026Claude Fable 5 (Max, Opus 4.8 Fallback) · 5 · Jun 9, 2026GPT-5.5 Instant (June 2026) · 5.5 · Jun 25, 2026Claude Sonnet 5 (Max) · 5 · Jun 30, 2026Grok 4.5 (High) · 4.5 · Jul 8, 2026GPT-5.6 Luna (Max) · 5.6 · Jul 9, 2026Gemini 3.5 Flash-Lite · 3.5 · Jul 21, 2026Gemini 3.6 Flash (High) · 3.6 · Jul 21, 2026Claude Opus 5 (Max) · 5 · Jul 24, 2026Grok 4.6 (High) · 4.6 · Aug 12, 2026Gemini 3.7 Flash (High) · 3.7 · Aug 13, 2026Claude Fable 5.1 (Max, Default Fallback) · 5.1 · Sep 1, 2026Gemini 3.8 Flash (High) · 3.8 · Sep 2, 2026GPT-6 Astra (Max) · 6 · Sep 3, 2026Grok 4.7 (Xhigh) · 4.7 · Sep 21, 2026Claude Opus 5.5 (Max, Default Fallback) · 5.5 · Sep 22, 2026GPT-6 Luna (Max) · 6 · Sep 22, 2026Claude Sonnet 5.5 (Max, Default Fallback) · 5.5 · Sep 28, 2026GPT-6.1 Sol (Max) · 6.1 · Sep 29, 2026Gemini 4 Argon (High) · 4 · Sep 30, 2026Claude Haiku 5.5 (Max) · 5.5 · Oct 7, 2026Claude Haiku 5.5Gemini 4 ArgonGPT-6.1 SolClaude Sonnet 5.5GPT-6 LunaClaude Opus 5.5Grok 4.7GPT-6 AstraGemini 3.8 FlashClaude Fable 5.1Gemini 3.7 FlashGrok 4.6Claude Opus 5Gemini 3.6 FlashGemini 3.5 Flash-LiteGPT-5.6 LunaGrok 4.5Claude Sonnet 5GPT-5.5 Instant (June2026)Claude Fable 5 (Max, Opus4.8 Fallback)Claude Opus 4.8Gemini 3.5 FlashGPT-5.5 Instant (May 2026)Grok 4.3GPT-5.5Claude Opus 4.7Grok 4.20 0309 v2GPT-5.4 miniGrok 4.20 0309GPT-5.4Gemini 3.1 Flash-LiteGemini 3.1 Pro PreviewClaude Sonnet 4.6GPT-5.3 CodexClaude Opus 4.6Gemini 3 Deep ThinkGemini 3 Flash PreviewGPT-5.2Claude Opus 4.5Grok 4.1 FastGemini 3 Pro PreviewGPT-5.1Claude 4.5 HaikuClaude 4.5 SonnetGemini 2.5 Flash-LitePreview (Sep '25)GPT-5 CodexGrok 4 FastGPT-5Claude 4.1 OpusGrok 4Gemini 2.5 Flash-Liteo3-proGemini 2.5 ProClaude 4 OpusGemini 2.5 FlashGemini 2.5 Pro Preview(May' 25)Gemini 2.5 Flash Previewo4-minio3GPT-4.1GPT-4o (March 2025,chatgpt-4o-latest)Gemini 2.5 Pro Preview(Mar' 25)o1-proGPT-4.5 (Preview)Gemini 2.0 Flash-Lite (Feb'25)Claude 3.7 SonnetGrok 3GPT-4o (ChatGPT)Gemini 2.0 Flash (Feb '25)o3-miniGemini 2.0 Flash ThinkingExperimental (Jan '25)Gemini 2.0 Flash ThinkingExperimental (Dec '24)GPT-4o mini Realtime (Dec'24)Grok 2 (Dec '24)Gemini 2.0 Flash(experimental)o1GPT-4o (Nov '24)Claude 3.5 HaikuGemini 1.5 Flash-8BGemini 1.5 Flash (Sep '24)o1-miniGPT-4o (Aug '24)GPT-4o miniClaude 3.5 Sonnet (June'24)Gemini 1.5 Pro (May '24)Gemini 1.5 Flash (May '24)GPT-4o (May '24)Grok-1Claude 3 HaikuGemini 1.0 ProClaude 2.1GPT-4 TurboClaude 2.0GPT-3.5 Turbo (0613)GPT-4GPT-3.5 TurboOpenAIAnthropicGooglexAIRelease datefranklineh.com/ridiclousbenchmark

Swipe the chart to see the full comparison.

Version Bench follows AI model releases using the number printed in the name. A model released later sits farther to the right. A model with a higher advertised version sits higher on the chart. GPT 6 beats GPT 5 under a scoring rule that requires absolutely no knowledge of either model's abilities.

Company colors help you follow releases from OpenAI, Google, Anthropic, and xAI. Use the release selector to include the full collection or focus on one company. Adjust the model names slider to show no names, all names, or a selection in between. Every dot stays visible, and downloaded charts use your chosen setting.

How we order AI model versions

Versions are compared one component at a time. The major number comes first, followed by the minor number and any further version component. That puts 6 above 5 and 5.10 above 5.9. Treating versions as ordinary decimal numbers would incorrectly turn 5.10 into 5.1, so this chart avoids that shortcut.

The vertical axis shows version order. The gap between two rows does not measure a gain in intelligence, speed, or accuracy. Lines connect releases within the same model family, making it easier to follow the sequence without implying that one company's numbering system matches another's.

Where the release dates come from

The timeline uses documented release or first preview dates from the model collection and matching official release announcements. A date on which Franklin first discovered a model is not a release date. Entries without both a usable date and a recognizable numeric version stay off the chart.

Configuration entries can describe the same family, version, and release date. Those shared points are combined so that a set of reasoning settings does not manufacture extra releases. The label identifies the catalogue entry retained for that point; hovering over the dot gives its date and version.

Reading the version arms race

Compare a family's releases over time, look for jumps in its numbering, or enjoy the fact that a larger number can make a model look victorious here. Model families choose their own numbering schemes. For a decision about which model to use, compare measured performance, cost, and the requirements of your task as well.

Explore the other ridiculous benchmarks or compare model performance in our AI benchmarks.