LLM Economics

OpenAI Cost Calculator

Estimate OpenAI API token spend with cached-input, Batch API, and long-context pricing rules.

Inputs

Choose the model tier you plan to use in production.

Average prompt size per API call — system prompt + context + user message. Typical ranges: simple chat 500–2,000 · RAG chatbot 2,000–8,000 · document analysis 10,000–100,000. A page of text ≈ 750 tokens.

Average response length per API call. Typical ranges: short answer 200–500 · paragraph 500–1,000 · long-form 1,000–4,000 tokens.

Total API calls per month across all users and jobs. 10 users × 3 sessions/day × 10 calls/session × 30 days = 9,000 requests/month.

Fraction of input tokens billed as cached reads. Cache writes and storage are not included. Use 0 if caching is disabled.

Estimated yearly cost
$96,774.20/year

((45,161.29 fresh input × $5 + 0 cached-read input × $0.5 + 19,354.84 output × $30) / 1M × 10,000 req) × 12

Decision Summary

Best move
Evaluate DeepSeek DeepSeek V4 Flash — potential $95,365.17/yr saving
Expected savings
$95,365.17 /yr (98.5% reduction)
Watch out
Validate quality, latency, and reliability on your own workload before changing production traffic.
Next step
Optimize with AI Cost Optimizer
Evaluate DeepSeek DeepSeek V4 Flash — potential $95,365.17/yr saving
At published token rates, DeepSeek DeepSeek V4 Flash handles the same token volume at $1,409.03/yr — 98.5% less than OpenAI GPT-5.6 Sol. This is a price-only comparison; validate quality, latency, tool support, and regional pricing before switching.
$95,365.17
saved / year (98.5%)
Prompt Cache not enabledEnable Prompt Cache

Enabling Prompt Cache with a 60% hit ratio could save $14,632.26/yr (15% reduction). Fix your system prompt at the start of every request and reuse it across calls.

$14,632.26
saved / year
Monthly cost
$8,064.52/month
Daily cost
$268.82/day
Cost per request
$0.806452
Model
GPT-5.6 Sol
Input rate
$5/MTok
Output rate
$30/MTok

Cost breakdown

ItemMonthlyYearly
Fresh input (45,161.29 tok/req × 10,000 req × $5/MTok)$2,258.06$27,096.77
Output spend (19,354.84 tok/req × 10,000 req × $30/MTok)$5,806.45$69,677.42
Total monthly spend$8,064.52$96,774.20

Comparison

OptionMonthlyYearly
OpenAI GPT-5.6 Solcurrent$8,064.52$96,774.20
Claude Claude Sonnet 5$2,838.71$34,064.52
Gemini Gemini 2.5 Pro$2,500.00$30,000.00
DeepSeek DeepSeek V4 Flashcheapest$117.42$1,409.03

Pricing sources

Last verified 2026-08-03 · OpenAI official pricing developers.openai.com/api/docs/models/compare

Industry Benchmark

Output price vs. peer average ($/1M tokens)Industry avg: 12.57 $/1M
You are at the 119th percentile

Pricing sources

Last verified 2026-08-03 · developers.openai.com/api/docs/models/compare developers.openai.com/api/docs/models/compare · platform.claude.com/docs/en/about-claude/pricing platform.claude.com/docs/en/about-claude/pricing · ai.google.dev/gemini-api/docs/pricing ai.google.dev/gemini-api/docs/pricing · api-docs.deepseek.com/quick_start/pricing api-docs.deepseek.com/quick_start/pricing

Trends & comparison

Trend

Comparison (monthly vs. yearly)

Official OpenAI rates used

Verified August 3, 2026: GPT-5 is $1.25/MTok input and $10 output; GPT-5 Mini is $0.25/$2; GPT-5 Nano is $0.05/$0.40. GPT-5.6 Sol is $5/$30, Terra is $2.50/$15, and Luna is $1/$6 before long-context pricing. Rates are stored in the shared pricing configuration and linked to OpenAI's official model documentation.

A reproducible workload example

A chatbot making 10,000 monthly calls with 2,000 input and 500 output tokens costs $75/month on GPT-5 before caching or Batch: 20 million input tokens at $1.25 plus 5 million output tokens at $10. The calculator exposes the same formula beside each result.

What the estimate includes

The estimate includes token usage for the selected model and the entered cached-read share. It excludes cache writes, cache storage, web search, containers, tool execution, fine-tuning, priority processing, taxes, credits, and enterprise agreements. Those items must be added from the applicable OpenAI price sheet or invoice.

Choosing a model economically

A lower token price does not prove equivalent quality. Compare candidate models on a representative evaluation set, then check latency, tool support, context limits, regional availability, and total non-token charges before moving production traffic.

Frequently asked questions

How is the monthly OpenAI API estimate calculated?▾

The calculator prices fresh input, cached-read input, and output tokens separately, then multiplies the per-request total by monthly request volume. Tool calls, web search, cache writes, storage, taxes, and contracted rates are not included.

When does GPT-5.6 long-context pricing apply?▾

For GPT-5.6 Sol, Terra, and Luna, requests above 272,000 input tokens use 2x input and 1.5x output rates for the entire request. The calculator applies that tier automatically.

How much can cached input save?▾

The result uses each model's published cached-input read rate only for the share entered as cache hits. Actual bills can also include cache-write or storage charges, so use provider billing data for final reconciliation.

When should I use the Batch API estimate?▾

Use it only for asynchronous jobs that qualify for OpenAI Batch pricing. The calculator applies the published 50% token-rate discount; it does not model service-specific charges.

Related calculators

Same cluster

Compare verified first-party API prices on the same standard workload → AI Model Price Leaderboard

OpenAI API Cost Calculator — GPT-5.6 & GPT-5 Token Pricing 2026