LLM Economics
17 models · Official sources checked 2026-08-03

AI Model Price Leaderboard

Compare first-party API token prices on one transparent workload. This leaderboard ranks cost only; it does not turn unrelated benchmark suites into an artificial quality score.

Verified API Cost Rankings

Standard workload

1,000 input + 500 output tokens per task. Monthly estimate assumes 10,000 users and 100 tasks per user.

#ModelSource
1

GPT-5 Nano

OpenAI
$0.250Official
2

DeepSeek V4 Flash

DeepSeek1M ctx
$0.280Official
3

Gemini 2.5 Flash-Lite

Gemini1.05M ctx
$0.300Official
4

DeepSeek V4 Pro

DeepSeek1M ctx
$0.870Official
5

GPT-5 Mini

OpenAI
$1.250Official
6

Gemini 2.5 Flash

Gemini1.05M ctx
$1.550Official
7

Claude Haiku 4.5

Claude200K ctx
$3.500Official
8

GPT-5.6 Luna

OpenAI1.05M ctx
$4.000Official
9

Gemini 3.6 Flash

Gemini1.05M ctx
$5.250Official
10

GPT-5

OpenAI
$6.250Official
11

Gemini 2.5 Pro

Gemini1.05M ctx
$6.250Official
12

Claude Sonnet 5

Claude1M ctx

Introductory pricing through 2026-08-31; standard pricing becomes $3/$15 per MTok on 2026-09-01.

$7.000Official
13

GPT-5.6 Terra

OpenAI1.05M ctx
$10.000Official
14

Claude Sonnet 4.6

Claude1M ctx
$10.500Official
15

Claude Opus 5

Claude1M ctx
$17.500Official
16

Claude Opus 4.7

Claude1M ctx
$17.500Official
17

GPT-5.6 Sol

OpenAI1.05M ctx
$20.000Official

Prices are first-party API rates, not reseller rates. Standard workload excludes tool-call fees, cache-write charges, regional premiums, storage, and long-context multipliers unless stated. Use each calculator for workload-specific pricing.

What This Ranking Can Tell You

It answers a narrow, reproducible question: what would the published token bill be for the same input and output volume? It is useful for budget screening and identifying models worth testing. It cannot tell you whether a cheaper model completes your task successfully, needs more reasoning tokens, makes more tool calls, or produces longer answers.

Important Pricing Conditions

Provider pricing has conditions that a compact table cannot fully express. Some OpenAI and Gemini models charge higher rates above long-context thresholds. Anthropic applies different rates for cache writes, cache reads, data residency, and fast mode. Tool calls, grounding, storage, and regional hosting can add separate charges. The workload calculators apply supported token tiers and show the exact formula used for the estimate.

Verification standard

Prices are published only when a first-party provider page confirms the model and token rates. Competitor databases such as Artificial Analysis and OpenRouter are used as anomaly checks, not as the source of record. Last checked 2026-08-03.

Frequently Asked Questions

How is the AI model price leaderboard ranked?

The default ranking uses a transparent standard workload: 1,000 input tokens plus 500 output tokens. Cost per 1,000 tasks is calculated directly from each provider's published per-million-token API rates.

Does the ranking include model quality?

No. This release intentionally ranks verified API cost only. Model quality depends on workload, prompting, reasoning effort, latency, and evaluation design. Run a representative eval before choosing a model based on price.

Are Batch API and prompt caching included?

The leaderboard uses standard uncached API rates so every model is compared on the same basis. The individual calculators can model cache hits, provider-supported batch discounts, and long-context price tiers.

Where do the prices come from?

Every row links directly to the provider's official pricing documentation and includes a verification date. Reseller and negotiated enterprise prices are not used.