Model Rankings
— Model rankings —

LLM rankings, side by side

Benchmarks, cost vs intelligence, context windows, and head-to-head comparisons across every major large language model.

Six lenses on the LLM field. Composite ranks for the headline answer; cost-curves and the heatmap for the nuance. Where you start depends on what you’re optimizing for.

Cost × Intelligence

Best value: high intelligence, low cost

Bubble size = context window. Best-value quadrant highlighted.

LLM Composite — Top 10

Average of AA’s same-scale intelligence indices (Overall and Coding) per model. Math is shown side-by-side at right but excluded from this rank — its 80–99 distribution would distort the average against models AA hasn’t scored on math yet.

#1GPT-5.5 (xhigh)
56.7—
#2GPT-5.5 (medium)
52.6—
#3DeepSeek V4 Pro 0813 (Reasoning, Max Effort)
52.4—
#4Claude Opus 4.7 (Non-reasoning, High Effort)
52.3—
#5DeepSeek V4 Flash 0731 (Reasoning, Max Effort)
51.7—
#6Gemini 3.1 Pro Preview
49.3—
#7GPT-5.5 (low)
45.8—
#8Muse Spark
45.0—
#9DeepSeek V4 Flash (Reasoning, High Effort)
44.8—
#10Claude Sonnet 4.6 (Non-reasoning, High Effort)
43.9—
Three indices: Overall · Coding · Math

The three intelligence indices Artificial Analysis publishes for every LLM, side-by-side for the top models. Missing values shown as —.

Overall
Coding
Math

Math shows as —when AA hasn’t run the underlying math benchmarks (AIME, MATH-500…) against a given variant — common for high-effort reasoning configurations.

GPT-5.5 (xhigh)
OpenAI
Overall
38
Coding
75
Math
—
GPT-5.5 (medium)
OpenAI
Overall
34
Coding
72
Math
—
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)
Overall
36
Coding
69
Math
—
Claude Opus 4.7 (Non-reasoning, High Effort)
Anthropic
Overall
31
Coding
74
Math
—
DeepSeek V4 Flash 0731 (Reasoning, Max Effort)
Overall
34
Coding
69
Math
—
Gemini 3.1 Pro Preview
Google DeepMind
Overall
30
Coding
69
Math
—
GPT-5.5 (low)
OpenAI
Overall
31
Coding
61
Math
—
Muse Spark
Overall
31
Coding
59
Math
—
Speed × Cost

Bubble size = intelligence. Bottom-right is the operational sweet spot: fast responses for low cost.

Benchmark Heatmap

Per-benchmark scores across the most-populated columns. Each column is normalized independently so different scales (MMLU 0-100 vs AIME 0-30) compare cleanly.

ModelOverall
n=12
Coding
n=12
GPQA
n=12
HLE
n=12
IFBench
n=12
LCR
n=12
SciCode
n=12
Tau Banking
n=12
Terminal-Bench (Hard)
n=12
Terminalbench V2 1
n=12
GPT-5.5 (xhigh)
OpenAI
38.474.90.940.460.760.840.560.390.610.84
DeepSeek V4 Flash (Reasoning, High Effort)
37.552.00.870.280.730.630.420.200.390.57
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)
36.068.80.930.410.710.800.510.400.420.79
DeepSeek V4 Flash 0731 (Reasoning, Max Effort)
34.369.10.910.390.790.800.500.390.360.79
GPT-5.5 (medium)
OpenAI
33.871.50.930.420.710.830.550.300.580.81
Muse Spark
31.358.60.880.410.760.780.520.200.450.62
Claude Opus 4.7 (Non-reasoning, High Effort)
Anthropic
30.973.60.890.330.440.760.500.350.550.83
GPT-5.5 (low)
OpenAI
30.760.90.910.330.640.810.520.250.520.66
Gemini 3.1 Pro Preview
Google DeepMind
29.768.80.940.470.770.820.590.210.540.74
GLM-5.1 (Reasoning)
26.155.80.870.300.760.740.450.140.430.62
MiMo-V2.5-Pro
26.060.20.870.360.800.800.510.100.430.65
Claude Sonnet 4.6 (Non-reasoning, High Effort)
Anthropic
24.763.00.800.130.410.680.500.340.460.71
Cyan border = best in column. Color intensity scaled per column. — = no data.
Cost Calculator

Estimate your monthly cost across models given your usage.

✓ Recommended: DeepSeek V4 Flash (Reasoning, High Effort) — $0.53/mo (within 5% of top intelligence)

Sources: artificialanalysis.ai, LMSYS Chatbot Arena, Stanford HELM, official model documentation. Composite scores updated when new benchmarks publish. Read our methodology →

Rankings data by Artificial Analysis. CSV imports cover supplementary benchmarks.