LLM Economics
30 models · Updated 2026-07-30

LLM Leaderboard

Frontier AI models ranked by composite benchmark score. Compare performance across coding, knowledge, reasoning, agentic tasks, and math — with real API pricing for cost-efficiency analysis.

Top Models

87

Claude Mythos 5

Anthropic

Best: Agentic (97.7)
$10/50 per M tokens
85

Claude Opus 5

Anthropic

Best: Knowledge (93.5)
$5/25 per M tokens
83

Kimi K3

Moonshot AI

Best: Agentic (94.7)
$3/15 per M tokens

Full Rankings

Model
1
Claude Mythos 5
Anthropic1M+
87
2
Claude Opus 5
Anthropic1M
85
3
Kimi K3
Moonshot AI1.05M
83
4
GPT-5.6 Sol
OpenAI1.05M
81
5
Claude Fable 5
Anthropic1M+
80
6
Claude Opus 4.8
Anthropic1M
78
7
Qwen3.7 Max
Alibaba1M
78
8
GPT-5.6 Terra
OpenAI1.05M
77
9
Claude Sonnet 5
Anthropic1M
76
10
Qwen3.7 Plus
Alibaba1M
74
11
Muse Spark 1.1
Meta1M
73
12
GPT-5.5
OpenAI1M
72
13
GPT-5.6 Luna
OpenAI1.05M
71
14
Gemini 3.6 Flash
Google1M
70
15
Claude Opus 4.7 (Adaptive)
Anthropic1M
69
16
Claude Opus 4.6
Anthropic1M
68
17
Grok 4.5
xAI500K
68
18
GPT-5.4
OpenAI1.05M
67
19
Claude Opus 4.7
Anthropic1M
66
20
GPT-5.3 Codex
OpenAI400K
64
21
GLM-5.1
Z.AI203K
63
22
GLM-5
Z.AI200K
63
23
Qwen3.6 Plus
Alibaba1M
63
24
Gemini 3.5 Flash-Lite
Google1M
61
25
Inkling
Thinking Machines Lab1M
60
26
MiMo-V2.5-Pro
Xiaomi1M
59
27
MiniMax M3
MiniMax1M
59
28
Muse Spark
Meta262K
57
29
Gemini 3 Pro
Google2M
57
30
GPT-5.4 nano
OpenAI400K
38

30 models shown · Score = weighted benchmark composite · Value = score per $1 of output tokens

How LLM Benchmarks Work in 2026

Modern LLM evaluation uses a combination of academic benchmarks, coding challenges, mathematical competitions, and real-world agentic tasks. No single benchmark captures overall model capability — the composite score here aggregates 27 weighted benchmarks across 8 categories to give a balanced view. Knowledge tests like HLE (Humanity's Last Exam) and MMLU-Pro measure factual breadth and graduate-level reasoning. Coding benchmarks like SWE-bench Verified and LiveCodeBench test the ability to fix real GitHub issues and solve competition problems. Mathematics evaluations include AIME, HMMT, and FrontierMath — ranging from olympiad-level to research-frontier difficulty.

The Frontier in July 2026

The top tier is dominated by Anthropic's Claude family and OpenAI's GPT-5.6 series. Claude Mythos 5 leads overall with a composite of 87 — its strength is agentic capability (97.7), making it the clear leader for autonomous tool-use and multi-step workflows. Claude Opus 5 follows at 85 with the most complete coverage across all categories including the highest reasoning score (92.5 on ARC-AGI-2 and LongBench v2). GPT-5.6 Sol excels in mathematics (97.3) and offers strong agentic performance (94.7), making it the top pick for STEM and research applications. Kimi K3 from Moonshot AI surprised the field with an 83 composite — particularly strong in multimodal tasks (89.6) and agentic benchmarks (94.7) — positioning Chinese labs as serious frontier competitors.

Cost vs Capability: Finding the Value Sweet Spot

Raw benchmark scores don't capture the full picture — cost per unit of capability matters for production deployments. GPT-5.6 Luna stands out as the best value play: scoring 71 overall with output pricing at just $1.20/M tokens, yielding a value metric of 55.47 score-per-dollar. Compare this to the flagship GPT-5.6 Sol at $30/M output (value: 2.71) — Luna delivers 20× more capability per dollar, though at a lower absolute performance level. For teams with fixed budgets, Claude Sonnet 5 at $10/M output (value: 6.45) offers a compelling middle ground: strong enough for most production tasks at 1/5 the cost of frontier models.

Choosing the Right Model for Your Use Case

The leaderboard reveals clear specialization patterns. For coding automation and software engineering agents, the Claude family leads uniformly (83.1 across Mythos, Opus 5, and Fable). For mathematical reasoning and research, GPT-5.6 Sol and Terra both score 97.3. For budget-conscious deployments needing broad capability, Gemini 3.6 Flash (score 70, $7.50/M output) and Grok 4.5 (score 68, $6/M output) offer frontier-adjacent performance at mid-tier pricing. Open-weight options like GLM-5.1 (63, $4.40/M) and MiniMax M3 (59, $1.20/M) provide self-hosting flexibility for privacy-sensitive workloads.

Key Trends to Watch

Several shifts define mid-2026: reasoning-enabled models now dominate the top 15 (only 4 non-reasoning models remain in the top 30). Context windows have standardized at 1M+ tokens for frontier models. The cost floor continues dropping — GPT-5.6 Luna and MiniMax M3 offer 70+ composite scores under $1.50/M output. And Chinese labs (Alibaba, Moonshot, MiniMax, Xiaomi, Z.AI) now hold 9 of the top 30 spots, up from 3 a year ago.

About This Leaderboard

Scoring Methodology

The composite score is a weighted average across key benchmarks: HLE (45%), MMLU-Pro (30%), GPQA (7%), SuperGPQA (7%), SimpleQA (11%) for knowledge; SWE-bench, LiveCodeBench, SciCode for coding; FrontierMath, AIME, HMMT for math; Terminal-Bench, BrowseComp, OSWorld for agentic tasks.

Use with LLM Economics

See how benchmark leaders perform in real cost scenarios with our calculators:

Frequently Asked Questions

What is the best LLM in 2026?

As of July 2026, Claude Mythos 5 by Anthropic leads the composite benchmark rankings with a score of 87, excelling in agentic tasks (97.7) and knowledge (93.8). Claude Opus 5 ranks second at 85, with the strongest reasoning score (92.5) among top models.

How are LLM benchmark scores calculated?

The composite score is a weighted average across 27 benchmarks in 8 categories: Knowledge (HLE 45%, MMLU-Pro 30%, GPQA 7%, SuperGPQA 7%, SimpleQA 11%), Coding (LiveCodeBench 38%, SWE-Rebench 20%, SWE-bench Verified 16%, SciCode 16%, SWE-bench Pro 10%), Mathematics (FrontierMath v2 30%, AIME 25%, HMMT 25%, USAMO 10%, FrontierMath Tier 4 10%), Reasoning (LongBench v2 38%, ARC-AGI-2 31%, MRCRv2 31%), and Agentic (Terminal-Bench 38%, OSWorld 34%, BrowseComp 28%).

Which LLM is cheapest per token?

GPT-5.6 Luna offers the best value among frontier models at $0.20 input / $1.20 output per million tokens with a composite score of 71 — yielding 55.47 score points per dollar of output. MiniMax M3 ($0.30/$1.20) and GPT-5.4 nano ($0.20/$1.25) also offer high value but with lower overall capability.

How does GPT-5.6 Sol compare to Claude Opus 5?

GPT-5.6 Sol (score 81) leads in math (97.3) and agentic tasks (94.7) but trails Claude Opus 5 (score 85) in coding (55 vs 83.1), knowledge (83.3 vs 93.5), and reasoning (94.3 vs 92.5). On price, both cost $5 input per million tokens but Sol is $30 vs Opus 5's $25 for output.

What is the best LLM for coding?

Claude Mythos 5, Claude Opus 5, and Claude Fable 5 all score 83.1 on coding benchmarks (SWE-bench Verified, LiveCodeBench, SciCode). Qwen3.7 Max follows at 80.4. Among budget options, Claude Sonnet 5 scores 61.4 at only $2/$10 per million tokens.

What is the best LLM for agentic tasks?

Claude Mythos 5 leads agentic benchmarks at 97.7, followed by GPT-5.6 Sol and Kimi K3 (both 94.7), Claude Fable 5 (92.9), and GPT-5.6 Terra (91.4). These scores are based on Terminal-Bench 2.0, BrowseComp, and OSWorld-Verified evaluations.

How often is this leaderboard updated?

The leaderboard is updated when significant new models are released or major benchmark results are published — typically 1-2 times per month. All scores are compiled from official provider system cards, independent evaluations (Artificial Analysis, Epoch AI), and public benchmark leaderboards.

What does the Value column mean?

Value = composite benchmark score divided by output price per million tokens. A higher value means you get more capability per dollar spent. GPT-5.6 Luna (55.47) and MiniMax M3 (57.3) lead this metric, meaning they offer the best performance-to-cost ratio among ranked models.