LLM Leaderboard
Frontier AI models ranked by composite benchmark score. Compare performance across coding, knowledge, reasoning, agentic tasks, and math — with real API pricing for cost-efficiency analysis.
Top Models
Claude Mythos 5
Anthropic
Claude Opus 5
Anthropic
Kimi K3
Moonshot AI
Full Rankings
| Model | ||
|---|---|---|
1 | Claude Mythos 5 Anthropic1M+ | 87 |
2 | Claude Opus 5 Anthropic1M | 85 |
3 | Kimi K3 Moonshot AI1.05M | 83 |
4 | GPT-5.6 Sol OpenAI1.05M | 81 |
5 | Claude Fable 5 Anthropic1M+ | 80 |
6 | Claude Opus 4.8 Anthropic1M | 78 |
7 | Qwen3.7 Max Alibaba1M | 78 |
8 | GPT-5.6 Terra OpenAI1.05M | 77 |
9 | Claude Sonnet 5 Anthropic1M | 76 |
10 | Qwen3.7 Plus Alibaba1M | 74 |
11 | Muse Spark 1.1 Meta1M | 73 |
12 | GPT-5.5 OpenAI1M | 72 |
13 | GPT-5.6 Luna OpenAI1.05M | 71 |
14 | Gemini 3.6 Flash Google1M | 70 |
15 | Claude Opus 4.7 (Adaptive) Anthropic1M | 69 |
16 | Claude Opus 4.6 Anthropic1M | 68 |
17 | Grok 4.5 xAI500K | 68 |
18 | GPT-5.4 OpenAI1.05M | 67 |
19 | Claude Opus 4.7 Anthropic1M | 66 |
20 | GPT-5.3 Codex OpenAI400K | 64 |
21 | GLM-5.1 Z.AI203K | 63 |
22 | GLM-5 Z.AI200K | 63 |
23 | Qwen3.6 Plus Alibaba1M | 63 |
24 | Gemini 3.5 Flash-Lite Google1M | 61 |
25 | Inkling Thinking Machines Lab1M | 60 |
26 | MiMo-V2.5-Pro Xiaomi1M | 59 |
27 | MiniMax M3 MiniMax1M | 59 |
28 | Muse Spark Meta262K | 57 |
29 | Gemini 3 Pro Google2M | 57 |
30 | GPT-5.4 nano OpenAI400K | 38 |
30 models shown · Score = weighted benchmark composite · Value = score per $1 of output tokens
How LLM Benchmarks Work in 2026
Modern LLM evaluation uses a combination of academic benchmarks, coding challenges, mathematical competitions, and real-world agentic tasks. No single benchmark captures overall model capability — the composite score here aggregates 27 weighted benchmarks across 8 categories to give a balanced view. Knowledge tests like HLE (Humanity's Last Exam) and MMLU-Pro measure factual breadth and graduate-level reasoning. Coding benchmarks like SWE-bench Verified and LiveCodeBench test the ability to fix real GitHub issues and solve competition problems. Mathematics evaluations include AIME, HMMT, and FrontierMath — ranging from olympiad-level to research-frontier difficulty.
The Frontier in July 2026
The top tier is dominated by Anthropic's Claude family and OpenAI's GPT-5.6 series. Claude Mythos 5 leads overall with a composite of 87 — its strength is agentic capability (97.7), making it the clear leader for autonomous tool-use and multi-step workflows. Claude Opus 5 follows at 85 with the most complete coverage across all categories including the highest reasoning score (92.5 on ARC-AGI-2 and LongBench v2). GPT-5.6 Sol excels in mathematics (97.3) and offers strong agentic performance (94.7), making it the top pick for STEM and research applications. Kimi K3 from Moonshot AI surprised the field with an 83 composite — particularly strong in multimodal tasks (89.6) and agentic benchmarks (94.7) — positioning Chinese labs as serious frontier competitors.
Cost vs Capability: Finding the Value Sweet Spot
Raw benchmark scores don't capture the full picture — cost per unit of capability matters for production deployments. GPT-5.6 Luna stands out as the best value play: scoring 71 overall with output pricing at just $1.20/M tokens, yielding a value metric of 55.47 score-per-dollar. Compare this to the flagship GPT-5.6 Sol at $30/M output (value: 2.71) — Luna delivers 20× more capability per dollar, though at a lower absolute performance level. For teams with fixed budgets, Claude Sonnet 5 at $10/M output (value: 6.45) offers a compelling middle ground: strong enough for most production tasks at 1/5 the cost of frontier models.
Choosing the Right Model for Your Use Case
The leaderboard reveals clear specialization patterns. For coding automation and software engineering agents, the Claude family leads uniformly (83.1 across Mythos, Opus 5, and Fable). For mathematical reasoning and research, GPT-5.6 Sol and Terra both score 97.3. For budget-conscious deployments needing broad capability, Gemini 3.6 Flash (score 70, $7.50/M output) and Grok 4.5 (score 68, $6/M output) offer frontier-adjacent performance at mid-tier pricing. Open-weight options like GLM-5.1 (63, $4.40/M) and MiniMax M3 (59, $1.20/M) provide self-hosting flexibility for privacy-sensitive workloads.
Key Trends to Watch
Several shifts define mid-2026: reasoning-enabled models now dominate the top 15 (only 4 non-reasoning models remain in the top 30). Context windows have standardized at 1M+ tokens for frontier models. The cost floor continues dropping — GPT-5.6 Luna and MiniMax M3 offer 70+ composite scores under $1.50/M output. And Chinese labs (Alibaba, Moonshot, MiniMax, Xiaomi, Z.AI) now hold 9 of the top 30 spots, up from 3 a year ago.
About This Leaderboard
Scoring Methodology
The composite score is a weighted average across key benchmarks: HLE (45%), MMLU-Pro (30%), GPQA (7%), SuperGPQA (7%), SimpleQA (11%) for knowledge; SWE-bench, LiveCodeBench, SciCode for coding; FrontierMath, AIME, HMMT for math; Terminal-Bench, BrowseComp, OSWorld for agentic tasks.
Use with LLM Economics
See how benchmark leaders perform in real cost scenarios with our calculators:
Frequently Asked Questions
What is the best LLM in 2026?
As of July 2026, Claude Mythos 5 by Anthropic leads the composite benchmark rankings with a score of 87, excelling in agentic tasks (97.7) and knowledge (93.8). Claude Opus 5 ranks second at 85, with the strongest reasoning score (92.5) among top models.
How are LLM benchmark scores calculated?
The composite score is a weighted average across 27 benchmarks in 8 categories: Knowledge (HLE 45%, MMLU-Pro 30%, GPQA 7%, SuperGPQA 7%, SimpleQA 11%), Coding (LiveCodeBench 38%, SWE-Rebench 20%, SWE-bench Verified 16%, SciCode 16%, SWE-bench Pro 10%), Mathematics (FrontierMath v2 30%, AIME 25%, HMMT 25%, USAMO 10%, FrontierMath Tier 4 10%), Reasoning (LongBench v2 38%, ARC-AGI-2 31%, MRCRv2 31%), and Agentic (Terminal-Bench 38%, OSWorld 34%, BrowseComp 28%).
Which LLM is cheapest per token?
GPT-5.6 Luna offers the best value among frontier models at $0.20 input / $1.20 output per million tokens with a composite score of 71 — yielding 55.47 score points per dollar of output. MiniMax M3 ($0.30/$1.20) and GPT-5.4 nano ($0.20/$1.25) also offer high value but with lower overall capability.
How does GPT-5.6 Sol compare to Claude Opus 5?
GPT-5.6 Sol (score 81) leads in math (97.3) and agentic tasks (94.7) but trails Claude Opus 5 (score 85) in coding (55 vs 83.1), knowledge (83.3 vs 93.5), and reasoning (94.3 vs 92.5). On price, both cost $5 input per million tokens but Sol is $30 vs Opus 5's $25 for output.
What is the best LLM for coding?
Claude Mythos 5, Claude Opus 5, and Claude Fable 5 all score 83.1 on coding benchmarks (SWE-bench Verified, LiveCodeBench, SciCode). Qwen3.7 Max follows at 80.4. Among budget options, Claude Sonnet 5 scores 61.4 at only $2/$10 per million tokens.
What is the best LLM for agentic tasks?
Claude Mythos 5 leads agentic benchmarks at 97.7, followed by GPT-5.6 Sol and Kimi K3 (both 94.7), Claude Fable 5 (92.9), and GPT-5.6 Terra (91.4). These scores are based on Terminal-Bench 2.0, BrowseComp, and OSWorld-Verified evaluations.
How often is this leaderboard updated?
The leaderboard is updated when significant new models are released or major benchmark results are published — typically 1-2 times per month. All scores are compiled from official provider system cards, independent evaluations (Artificial Analysis, Epoch AI), and public benchmark leaderboards.
What does the Value column mean?
Value = composite benchmark score divided by output price per million tokens. A higher value means you get more capability per dollar spent. GPT-5.6 Luna (55.47) and MiniMax M3 (57.3) lead this metric, meaning they offer the best performance-to-cost ratio among ranked models.