← All ridiculous benchmarks

Calendar Bench

Calendar BenchGPT 6: 301. 18 active variants; 18 with release dates.; Claude 5.5: 285. 15 active variants; 15 with release dates.; GPT 6.1: 145. 5 active variants; 5 with release dates.; Grok 4.7: 63. 3 active variants; 3 with release dates.; Grok 4.6: 48. 4 active variants; 4 with release dates.; Gemini 4: 30. 1 active variant; 1 with release dates.; Gemini 3.8: 6. 3 active variants; 3 with release dates.; Claude 5.1: 5. 5 active variants; 5 with release dates.Calendar BenchThe 31st beats the 1st. More variants bring more calendar.Calendar points061122183244305GPT 6: 301 Calendar points. 18 active variants; 18 with release dates.301GPT 6Claude 5.5: 285 Calendar points. 15 active variants; 15 with release dates.285Claude 5.5GPT 6.1: 145 Calendar points. 5 active variants; 5 with release dates.145GPT 6.1Grok 4.7: 63 Calendar points. 3 active variants; 3 with release dates.63Grok 4.7Grok 4.6: 48 Calendar points. 4 active variants; 4 with release dates.48Grok 4.6Gemini 4: 30 Calendar points. 1 active variant; 1 with release dates.30Gemini 4Gemini 3.8: 6 Calendar points. 3 active variants; 3 with release dates.6Gemini 3.8Claude 5.1: 5 Calendar points. 5 active variants; 5 with release dates.5Claude 5.1Model generationfranklineh.com/ridiclousbenchmark

Swipe the chart to see the full comparison.

Calendar Bench turns release dates into a wonderfully unhelpful leaderboard. A variant released on the 1st earns one point. A variant released on the 31st earns 31. We add the release-day numbers of the active variants in a model generation and show the combined total as one bar.

This rewards two things: launching late in the month and having plenty of variants. Month and year do not increase the score. A launch on January 31 earns the same points as one on October 31, regardless of what the model can actually do.

Adding release days across model variants

We read each variant's documented release date and take its day of the month in UTC. Three variants released on the 3rd, 17th, and 31st produce a generation total of 51 points. If several distinct configurations share a release date, each contributes that day's number to the total.

Only real release or first preview dates count. We use dates in the model collection and matching release announcements already tracked by Franklin. The date on which we discovered a model cannot stand in for its launch date. Invalid dates and dates in the future are excluded.

Grouping entries into one model generation

Calendar Bench uses the same grouping as Variant Bench: creator, model family, and advertised generation. GPT 6 entries share one bar, while GPT 6.1 has another. Product tiers and reasoning settings count separately when the catalogue identifies them as distinct variants.

Retired entries are excluded, and duplicate source IDs or legacy aliases count once. Choosing any variant in the picker includes its full active generation. Choosing two variants from that generation still produces one combined bar, so you can compare whole groups without selecting every configuration yourself.

Reading incomplete dates and daily updates

Hover over a bar to see the number of active variants and how many have usable release dates. A partial total includes only dated variants. A generation with no usable dates gets an Undated marker rather than an invented zero. Missing dates can therefore affect its position in the leaderboard.

The nightly collection can add variants, fill in dates, or retire entries. Those changes update the total when the page reads the refreshed catalogue. The default covers leading generations from OpenAI, Google, Anthropic, and xAI; the picker includes the wider collection. Export a chart with Franklin's credit when a calendar-based victory deserves to be shared.

Explore the other ridiculous benchmarks or compare model performance in our AI benchmarks.