Scrabble Bench scores AI model names using the standard English Scrabble tile values. Each letter contributes its tile score, and the chart adds those values to produce the model's total. A short name with expensive letters can compete with a much longer one.
This gives the model naming department a rare chance to influence a leaderboard directly. A well-placed Q or Z contributes ten points before the model answers a single question.
How the Scrabble score is calculated
A, E, I, L, N, O, R, S, T, and U are worth one point each. D and G are worth two; B, C, M, and P are worth three. F, H, V, W, and Y contribute four. K is worth five, J and X eight, and Q and Z ten.
We add the values of the A-Z letters in the displayed product name. Capitalization makes no difference. Spaces, numbers, punctuation, and letters outside that alphabet contribute zero. There are no double-word squares, triple-letter squares, or bonuses for using seven tiles.
Why name length and tile value disagree
Gemini 3.8 Flash scores 20 Scrabble points. Gemini contributes nine and Flash contributes eleven; the version number and spaces add nothing. Character Bench gives the same displayed name 16 characters because it counts those spaces and punctuation as well.
A model can therefore move up one chart and down the other without any change in its performance. Adding more digits increases the character count but leaves the Scrabble score alone. Adding an X adds one character and eight Scrabble points. Both calculations use the same cleaned product name, with evaluation settings removed.
Compare the names in your collection
Start with the flagship selection or choose up to 16 models from the full collection. The bars sort by total tile value, and you can download the result with Franklin's credit included. Use this leaderboard to settle a naming argument or make a chart worth sharing; use task-specific evaluations when you need to choose an AI model for work.
Explore the other ridiculous benchmarks or compare model performance in our AI benchmarks.