Quality, cost, speed, and efficiency
Model benchmarks
There are too many models and every provider says theirs is the best. This page puts quality, cost, speed, latency and token use in one place so I can compare them without opening fifteen tabs.
The Pareto frontier keeps the models that no other model beats across all your selected metrics. Pick X and Y for a normal chart, then add Z if you want the version you can spin around.
370 model records. Refreshed every 48 hours. Checked .
Current data
A few quick picks
These come from the latest data using the rule on each card, so the names can change after a refresh.
Muse Spark 1.3 (Max)
Meta Language
- Intelligence
- 48.1
- Output speed
- 159 t/s
Highest speed among models scoring at least 80% of the top intelligence score.
MiMo-V2.6-Pro
Xiaomi Language
- Intelligence
- 46.3
- Output speed
- 42 t/s
Highest intelligence score among open-weight models.
GPT-6 Luna (Low)
OpenAI Language
- Intelligence
- 21.5
- Cost per task
- $0.0045
Lowest measured task cost among models with an intelligence score.
GPT Image 2.5 Sunburst (max)
OpenAI Image
- Quality Elo
- 1197
- Price
- $211
Highest Image Arena quality score.
Grok Voice Think Fast 2.0 High
SpaceXAI Voice
- Quality
- 82.9%
- First audio
- 0.7s
Lowest latency among voice models scoring at least 90% of the top quality score.
Wan 3.0
Alibaba Video
- Quality Elo
- 1159
- Price
- $12
Highest Video Arena quality score.
Metric definitions
Four measures used in this explorer
Token efficiency
Output tokens per benchmark task measures how many tokens a language model uses to reach its score. Fewer tokens can matter when a subscription, gateway, or internal system has a fixed token budget.
Cost efficiency
Cost per task combines token use with input, cache, reasoning, and output prices. It gives a direct measure for metered API workloads. Token use and cost can produce different rankings.
Latency and throughput
Time to first output measures the initial wait. Output speed measures how quickly the full response arrives. The explorer includes both metrics.
Beyond language models
Image, voice, and video models have separate metric sets. These include Arena quality, media pricing, time to first audio, and rating evidence. Select a category to view its metrics.
How the frontier is computed
A model is on the frontier when no other visible model matches or exceeds every selected metric and improves at least one metric. Filters and axis changes recompute the set in your browser. The 2D line and 3D surface connect measured frontier points as a visual guide.
Metric coverage
Some measurements are unavailable for certain models. The caption reports how many records the current axis choice excludes. Detailed token data covers fewer language models than pricing, speed, and latency data.
Source and refresh
Data comes from public Artificial Analysis pages. A GitHub Action checks daily and refreshes the data when the current dataset is at least 47 hours old. This produces a refresh roughly every 48 hours. The previous dataset remains available if validation fails.