WhatLLM.org

Tool snapshot
Best fitCoding · Automation · Research
In one lineWhatLLM.org compares language models by quality, speed, token pricing, and context length, with task-specific rankings and side-by-side comparisons.

WhatLLM.org is a comparison site for people choosing a language model or provider. Its official comparison homepage combines benchmark data, token pricing, and output throughput in a searchable model database. It also provides focused rankings for coding, agent workflows, local deployment, and long-context work.

Comparing model tradeoffs

WhatLLM compares models on quality, speed, price, and context length. Its quality table uses the Artificial Analysis Intelligence Index, a composite capability score. The publisher explains that the evaluation set and weights depend on the methodology version, and recommends comparing scores from the same version before using task-level results for a particular workload.

Speed is shown in output tokens per second. Price uses per-million-token figures for input and output, alongside blended costs. Context length represents the maximum tokens a model can process in a single request. These separate dimensions let you compare a shortlist around the requirements of a task, rather than relying on a single ranking number.

Building a shortlist for a task

The Compare page lets you select two to four models for a side-by-side breakdown. WhatLLM also describes an LLM Selector that asks about the intended use, such as coding, analysis, creative writing, or agentic workflows, and recommends models ranked by fit.

Its ranking library includes local coding and Ollama lists as well as broader open-weight and agent rankings. The local-model category maps choices to hardware constraints, including Mac, RTX, workstation, and server use. Coding rankings refer to benchmarks such as LiveCodeBench, Terminal-Bench, and SciCode. The site also publishes analysis on releases, benchmark trends, and deployment costs.

Understanding the evidence and refresh dates

WhatLLM says its benchmark and quality data comes from Artificial Analysis. It does not run its own benchmarks. Provider pricing comes from the same source, while reviewed model guides cite official documentation.

The homepage says benchmark responses may be cached for up to 24 hours. It distinguishes data-fetch and snapshot dates from the date a benchmark was measured, and separates fetch status from editorial review dates. Those distinctions matter when comparing a newly released model with an older benchmark result. Rankings and model prices are live data, so use the current database for a purchasing decision rather than treating a captured leaderboard position as permanent.