Model Rankings
— Model rankings —

LLM rankings, side by side

Benchmarks, cost vs intelligence, context windows, and head-to-head comparisons across every major large language model.

Six lenses on the LLM field. Composite ranks for the headline answer; cost-curves and the heatmap for the nuance. Where you start depends on what you’re optimizing for.

Cost × Intelligence

Best value: high intelligence, low cost

Bubble size = context window. Best-value quadrant highlighted.

LLM Composite — Top 10

Average of AA’s same-scale intelligence indices (Overall and Coding) per model. Math is shown side-by-side at right but excluded from this rank — its 80–99 distribution would distort the average against models AA hasn’t scored on math yet.

#1GPT-5.5 (xhigh)
65.6
#2GPT-5.5 (medium)
61.5
#3DeepSeek V4 Flash 0731 (Reasoning, Max Effort)
60.4
#4Claude Opus 4.7 (Non-reasoning, High Effort)
58.8
#5Gemini 3.1 Pro Preview
58.3
#6Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)
55.7
#7GPT-5.5 (low)
52.7
#8MiMo-V2.5-Pro
51.5
#9Muse Spark
51.5
#10DeepSeek V4 Pro (Reasoning, High Effort)
51.2
Three indices: Overall · Coding · Math

The three intelligence indices Artificial Analysis publishes for every LLM, side-by-side for the top models. Missing values shown as —.

Overall
Coding
Math

Math shows as when AA hasn’t run the underlying math benchmarks (AIME, MATH-500…) against a given variant — common for high-effort reasoning configurations.

GPT-5.5 (xhigh)
OpenAI
Overall
56
Coding
75
Math
GPT-5.5 (medium)
OpenAI
Overall
51
Coding
72
Math
DeepSeek V4 Flash 0731 (Reasoning, Max Effort)
Overall
52
Coding
69
Math
Claude Opus 4.7 (Non-reasoning, High Effort)
Anthropic
Overall
44
Coding
74
Math
Gemini 3.1 Pro Preview
Google DeepMind
Overall
48
Coding
69
Math
Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)
Anthropic
Overall
48
Coding
63
Math
GPT-5.5 (low)
OpenAI
Overall
45
Coding
61
Math
MiMo-V2.5-Pro
Overall
43
Coding
60
Math
Speed × Cost

Bubble size = intelligence. Bottom-right is the operational sweet spot: fast responses for low cost.

Benchmark Heatmap

Per-benchmark scores across the most-populated columns. Each column is normalized independently so different scales (MMLU 0-100 vs AIME 0-30) compare cleanly.

ModelOverall
n=12
Coding
n=12
GPQA
n=12
HLE
n=12
IFBench
n=12
LCR
n=12
SciCode
n=12
Tau Banking
n=12
Terminal-Bench (Hard)
n=12
Terminalbench V2 1
n=12
GPT-5.5 (xhigh)
OpenAI
56.374.90.940.460.760.790.560.390.610.84
DeepSeek V4 Flash 0731 (Reasoning, Max Effort)
51.869.10.910.390.790.740.500.390.360.79
GPT-5.5 (medium)
OpenAI
51.471.50.930.420.710.770.540.300.580.81
Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)
Anthropic
48.463.00.880.340.570.740.470.340.530.71
Gemini 3.1 Pro Preview
Google DeepMind
47.768.80.940.470.770.790.590.210.540.74
GPT-5.5 (low)
OpenAI
44.560.90.910.330.640.780.520.250.520.66
Muse Spark
44.358.60.880.410.760.770.520.200.450.62
Claude Opus 4.7 (Non-reasoning, High Effort)
Anthropic
43.973.60.890.330.440.720.500.350.550.83
DeepSeek V4 Pro (Reasoning, High Effort)
43.758.70.910.350.710.670.460.260.420.65
MiMo-V2.5-Pro
42.960.20.870.360.800.780.500.100.430.65
GLM-5.1 (Reasoning)
41.055.80.870.300.760.680.440.140.430.62
GPT-5.4 mini (xhigh)
OpenAI
40.956.10.880.280.730.730.500.260.520.59
Cyan border = best in column. Color intensity scaled per column. — = no data.
Cost Calculator

Estimate your monthly cost across models given your usage.

✓ Recommended: GPT-5.5 (xhigh) — $33.75/mo (within 5% of top intelligence)

Sources: artificialanalysis.ai, LMSYS Chatbot Arena, Stanford HELM, official model documentation. Composite scores updated when new benchmarks publish. Read our methodology →

Rankings data by Artificial Analysis. CSV imports cover supplementary benchmarks.