
LLM rankings, side by side
Benchmarks, cost vs intelligence, context windows, and head-to-head comparisons across every major large language model.
Six lenses on the LLM field. Composite ranks for the headline answer; cost-curves and the heatmap for the nuance. Where you start depends on what you’re optimizing for.
Best value: high intelligence, low cost
Bubble size = context window. Best-value quadrant highlighted.
Average of AA’s same-scale intelligence indices (Overall and Coding) per model. Math is shown side-by-side at right but excluded from this rank — its 80–99 distribution would distort the average against models AA hasn’t scored on math yet.
The three intelligence indices Artificial Analysis publishes for every LLM, side-by-side for the top models. Missing values shown as —.
Math shows as —when AA hasn’t run the underlying math benchmarks (AIME, MATH-500…) against a given variant — common for high-effort reasoning configurations.
Bubble size = intelligence. Bottom-right is the operational sweet spot: fast responses for low cost.
Per-benchmark scores across the most-populated columns. Each column is normalized independently so different scales (MMLU 0-100 vs AIME 0-30) compare cleanly.
| Model | Overall n=12 | Coding n=12 | GPQA n=12 | HLE n=12 | IFBench n=12 | LCR n=12 | SciCode n=12 | Tau Banking n=12 | Terminal-Bench (Hard) n=12 | Terminalbench V2 1 n=12 |
|---|---|---|---|---|---|---|---|---|---|---|
GPT-5.5 (xhigh) OpenAI | 56.3 | 74.9 | 0.94 | 0.46 | 0.76 | 0.79 | 0.56 | 0.39 | 0.61 | 0.84 |
DeepSeek V4 Flash 0731 (Reasoning, Max Effort) | 51.8 | 69.1 | 0.91 | 0.39 | 0.79 | 0.74 | 0.50 | 0.39 | 0.36 | 0.79 |
GPT-5.5 (medium) OpenAI | 51.4 | 71.5 | 0.93 | 0.42 | 0.71 | 0.77 | 0.54 | 0.30 | 0.58 | 0.81 |
Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort) Anthropic | 48.4 | 63.0 | 0.88 | 0.34 | 0.57 | 0.74 | 0.47 | 0.34 | 0.53 | 0.71 |
Gemini 3.1 Pro Preview Google DeepMind | 47.7 | 68.8 | 0.94 | 0.47 | 0.77 | 0.79 | 0.59 | 0.21 | 0.54 | 0.74 |
GPT-5.5 (low) OpenAI | 44.5 | 60.9 | 0.91 | 0.33 | 0.64 | 0.78 | 0.52 | 0.25 | 0.52 | 0.66 |
Muse Spark | 44.3 | 58.6 | 0.88 | 0.41 | 0.76 | 0.77 | 0.52 | 0.20 | 0.45 | 0.62 |
Claude Opus 4.7 (Non-reasoning, High Effort) Anthropic | 43.9 | 73.6 | 0.89 | 0.33 | 0.44 | 0.72 | 0.50 | 0.35 | 0.55 | 0.83 |
DeepSeek V4 Pro (Reasoning, High Effort) | 43.7 | 58.7 | 0.91 | 0.35 | 0.71 | 0.67 | 0.46 | 0.26 | 0.42 | 0.65 |
MiMo-V2.5-Pro | 42.9 | 60.2 | 0.87 | 0.36 | 0.80 | 0.78 | 0.50 | 0.10 | 0.43 | 0.65 |
GLM-5.1 (Reasoning) | 41.0 | 55.8 | 0.87 | 0.30 | 0.76 | 0.68 | 0.44 | 0.14 | 0.43 | 0.62 |
GPT-5.4 mini (xhigh) OpenAI | 40.9 | 56.1 | 0.88 | 0.28 | 0.73 | 0.73 | 0.50 | 0.26 | 0.52 | 0.59 |
Estimate your monthly cost across models given your usage.
Sources: artificialanalysis.ai, LMSYS Chatbot Arena, Stanford HELM, official model documentation. Composite scores updated when new benchmarks publish. Read our methodology →
Rankings data by Artificial Analysis. CSV imports cover supplementary benchmarks.