Calendar Bench turns release dates into a wonderfully unhelpful leaderboard. A variant released on the 1st earns one point. A variant released on the 31st earns 31. We add the release-day numbers of the active variants in a model generation and show the combined total as one bar.
This rewards two things: launching late in the month and having plenty of variants. Month and year do not increase the score. A launch on January 31 earns the same points as one on October 31, regardless of what the model can actually do.
Adding release days across model variants
We read each variant's documented release date and take its day of the month in UTC. Three variants released on the 3rd, 17th, and 31st produce a generation total of 51 points. If several distinct configurations share a release date, each contributes that day's number to the total.
Only real release or first preview dates count. We use dates in the model collection and matching release announcements already tracked by Franklin. The date on which we discovered a model cannot stand in for its launch date. Invalid dates and dates in the future are excluded.
Grouping entries into one model generation
Calendar Bench uses the same grouping as Variant Bench: creator, model family, and advertised generation. GPT 6 entries share one bar, while GPT 6.1 has another. Product tiers and reasoning settings count separately when the catalogue identifies them as distinct variants.
Retired entries are excluded, and duplicate source IDs or legacy aliases count once. Choosing any variant in the picker includes its full active generation. Choosing two variants from that generation still produces one combined bar, so you can compare whole groups without selecting every configuration yourself.
Reading incomplete dates and daily updates
Hover over a bar to see the number of active variants and how many have usable release dates. A partial total includes only dated variants. A generation with no usable dates gets an Undated marker rather than an invented zero. Missing dates can therefore affect its position in the leaderboard.
The nightly collection can add variants, fill in dates, or retire entries. Those changes update the total when the page reads the refreshed catalogue. The default covers leading generations from OpenAI, Google, Anthropic, and xAI; the picker includes the wider collection. Export a chart with Franklin's credit when a calendar-based victory deserves to be shared.
Explore the other ridiculous benchmarks or compare model performance in our AI benchmarks.