Version Bench follows AI model releases using the number printed in the name. A model released later sits farther to the right. A model with a higher advertised version sits higher on the chart. GPT 6 beats GPT 5 under a scoring rule that requires absolutely no knowledge of either model's abilities.
Company colors help you follow releases from OpenAI, Google, Anthropic, and xAI. Use the release selector to include the full collection or focus on one company. Adjust the model names slider to show no names, all names, or a selection in between. Every dot stays visible, and downloaded charts use your chosen setting.
How we order AI model versions
Versions are compared one component at a time. The major number comes first, followed by the minor number and any further version component. That puts 6 above 5 and 5.10 above 5.9. Treating versions as ordinary decimal numbers would incorrectly turn 5.10 into 5.1, so this chart avoids that shortcut.
The vertical axis shows version order. The gap between two rows does not measure a gain in intelligence, speed, or accuracy. Lines connect releases within the same model family, making it easier to follow the sequence without implying that one company's numbering system matches another's.
Where the release dates come from
The timeline uses documented release or first preview dates from the model collection and matching official release announcements. A date on which Franklin first discovered a model is not a release date. Entries without both a usable date and a recognizable numeric version stay off the chart.
Configuration entries can describe the same family, version, and release date. Those shared points are combined so that a set of reasoning settings does not manufacture extra releases. The label identifies the catalogue entry retained for that point; hovering over the dot gives its date and version.
Reading the version arms race
Compare a family's releases over time, look for jumps in its numbering, or enjoy the fact that a larger number can make a model look victorious here. Model families choose their own numbering schemes. For a decision about which model to use, compare measured performance, cost, and the requirements of your task as well.
Explore the other ridiculous benchmarks or compare model performance in our AI benchmarks.