Quality, cost, speed, and efficiency

Model benchmarks

There are too many models and every provider says theirs is the best. This page puts quality, cost, speed, latency and token use in one place so I can compare them without opening fifteen tabs.

The Pareto frontier keeps the models that no other model beats across all your selected metrics. Pick X and Y for a normal chart, then add Z if you want the version you can spin around.

370 model records. Refreshed every 48 hours. Checked .

Current data

A few quick picks

These come from the latest data using the rule on each card, so the names can change after a refresh.

Fast and capable

Muse Spark 1.3 (Max)

Meta Language

Intelligence
48.1
Output speed
159 t/s

Highest speed among models scoring at least 80% of the top intelligence score.

Open weights

MiMo-V2.6-Pro

Xiaomi Language

Intelligence
46.3
Output speed
42 t/s

Highest intelligence score among open-weight models.

Lowest task cost

GPT-6 Luna (Low)

OpenAI Language

Intelligence
21.5
Cost per task
$0.0045

Lowest measured task cost among models with an intelligence score.

Image generation

GPT Image 2.5 Sunburst (max)

OpenAI Image

Quality Elo
1197
Price
$211

Highest Image Arena quality score.

Realtime voice

Grok Voice Think Fast 2.0 High

SpaceXAI Voice

Quality
82.9%
First audio
0.7s

Lowest latency among voice models scoring at least 90% of the top quality score.

Video generation

Wan 3.0

Alibaba Video

Quality Elo
1159
Price
$12

Highest Video Arena quality score.

Metric definitions

Four measures used in this explorer

Token efficiency

Output tokens per benchmark task measures how many tokens a language model uses to reach its score. Fewer tokens can matter when a subscription, gateway, or internal system has a fixed token budget.

Cost efficiency

Cost per task combines token use with input, cache, reasoning, and output prices. It gives a direct measure for metered API workloads. Token use and cost can produce different rankings.

Latency and throughput

Time to first output measures the initial wait. Output speed measures how quickly the full response arrives. The explorer includes both metrics.

Beyond language models

Image, voice, and video models have separate metric sets. These include Arena quality, media pricing, time to first audio, and rating evidence. Select a category to view its metrics.

How the frontier is computed

A model is on the frontier when no other visible model matches or exceeds every selected metric and improves at least one metric. Filters and axis changes recompute the set in your browser. The 2D line and 3D surface connect measured frontier points as a visual guide.

Metric coverage

Some measurements are unavailable for certain models. The caption reports how many records the current axis choice excludes. Detailed token data covers fewer language models than pricing, speed, and latency data.

Source and refresh

Data comes from public Artificial Analysis pages. A GitHub Action checks daily and refreshes the data when the current dataset is at least 47 hours old. This produces a refresh roughly every 48 hours. The previous dataset remains available if validation fails.