AI-generated content. Everything on this page was written by an automated AI editorial system and published without prior human review. How this works ›

Calculator · rates checked Oct 9, 2026

LLM API cost calculator

What the same workload costs on Claude, GPT-6, Gemini, Grok, DeepSeek and Mistral, with prompt caching, batch discounts, long-prompt surcharges and each vendor's tokenizer applied. The last one is the one most calculators skip, and it can move a bill by half.

Start from

5,000 conversations a day; a long system prompt reused from cache.

Models 15 selected

· ·

Anthropic

OpenAI

Google

xAI

Mistral

DeepSeek

Monthly cost, cheapest first · 30-day month, USD

Cheapest of the 15 selected: GPT-6 Luna at $51 a month. Claude Fable 5.1 costs 143× as much ($7,233).

Same text, different token counts. Your sizes are in OpenAI tokens. For the same text, Claude 4.7 and later about 49% more; DeepSeek about 20% more. Untick “Adjust for each model’s tokenizer” to see the naive comparison.

  1. GPT-6 LunaOpenAI
    $51/mo
  2. Claude Haiku 5.5Anthropic
    $75/mo
  3. DeepSeek V4.1 FlashDeepSeek
    $93/mo
    Includes weekday peak hours (2×)
  4. Gemini 3.5 Flash-LiteGoogle
    $212/mo
  5. DeepSeek V4 ProDeepSeek
    $353/mo
    Includes weekday peak hours (2×)
  6. Gemini 3.8 FlashGoogle
    $380/mo
    ▲ $760/mo after Dec 31
  7. Mistral Large 4Mistral
    $533/mo
  8. Grok 4.7xAI
    $855/mo
  9. GPT-6.1 SolOpenAI
    $987/mo
  10. Mistral Medium 3.5Mistral
    $1,125/mo
    No cache discount
  11. Gemini 3.1 Pro (preview)Google
    $1,134/mo
    Preview
  12. Claude Sonnet 5.5Anthropic
    $1,467/mo
  13. Claude Opus 5.5Anthropic
    $2,933/mo
  14. GPT-6 AstraOpenAI
    $5,070/mo
  15. Claude Fable 5.1Anthropic
    $7,233/mo
Show the numbers as a table
ModelTokens inTokens outPer requestPer 1,000Per month
GPT-6 LunaOpenAI · gpt-6-luna 3,000400 $0.0003$0.34 $51
Claude Haiku 5.5Anthropic · claude-haiku-5-5 4,458594 $0.0005$0.50 $75
DeepSeek V4.1 FlashDeepSeek · deepseek-flash 3,600480 $0.0006$0.62 $93
Gemini 3.5 Flash-LiteGoogle · gemini-3.5-flash-lite 3,000400 $0.0014$1.41 $212
DeepSeek V4 ProDeepSeek · deepseek-v4-pro 3,600480 $0.0024$2.35 $353
Gemini 3.8 FlashGoogle · gemini-3.8-flash 3,000400 $0.0025$2.54 $380
Mistral Large 4Mistral · mistral-large-4 3,000400 $0.0036$3.56 $533
Grok 4.7xAI · grok-4.7 3,000400 $0.0057$5.70 $855
GPT-6.1 SolOpenAI · gpt-6.1-sol 3,000400 $0.0066$6.58 $987
Mistral Medium 3.5Mistral · mistral-medium-3.5 3,000400 $0.0075$7.50 $1,125
Gemini 3.1 Pro (preview)Google · gemini-3.1-pro-preview 3,000400 $0.0076$7.56 $1,134
Claude Sonnet 5.5Anthropic · claude-sonnet-5-5 4,458594 $0.0098$9.78 $1,467
Claude Opus 5.5Anthropic · claude-opus-5-5 4,458594 $0.020$20 $2,933
GPT-6 AstraOpenAI · gpt-6-astra 3,000400 $0.034$34 $5,070
Claude Fable 5.1Anthropic · claude-fable-5-1 4,458594 $0.048$48 $7,233

Worked examples

The current lineup from each vendor on three common workloads, sizes counted in OpenAI tokens and converted per model. Load any of them into the calculator above to change the numbers.

Support chatbot

5,000 conversations a day; a long system prompt reused from cache. 3,000 in / 400 out per request, 60% cached.

ModelPer month
GPT-6 Luna $51
Claude Haiku 5.5 $75
DeepSeek V4.1 Flash $93
Gemini 3.5 Flash-Lite $212
DeepSeek V4 Pro $353
Gemini 3.8 Flash $380
Mistral Large 4 $533
Grok 4.7 $855
GPT-6.1 Sol $987
Mistral Medium 3.5 $1,125
Gemini 3.1 Pro (preview) $1,134
Claude Sonnet 5.5 $1,467
Claude Opus 5.5 $2,933
GPT-6 Astra $5,070
Claude Fable 5.1 $7,233

Coding agent

800 agent turns a day over a large, mostly cached repository context. 60K in / 2,500 out per request, 85% cached.

ModelPer month
GPT-6 Luna $64
Claude Haiku 5.5 $95
DeepSeek V4.1 Flash $105
Gemini 3.5 Flash-Lite $252
DeepSeek V4 Pro $418
Gemini 3.8 Flash $479
Mistral Large 4 $716
GPT-6.1 Sol $1,154
Gemini 3.1 Pro (preview) $1,397
Grok 4.7 $1,404
Claude Sonnet 5.5 $1,715
Mistral Medium 3.5 $2,610
Claude Opus 5.5 $3,431
GPT-6 Astra $6,384
Claude Fable 5.1 $8,122

Long documents

50 contracts a day at 300,000 tokens each — past most long-context thresholds. 300K in / 3,000 out per request, 50% cached.

ModelPer month
GPT-6 Luna $53
DeepSeek V4.1 Flash $54
Gemini 3.5 Flash-Lite $85
Claude Haiku 5.5 $201
Gemini 3.8 Flash $203
DeepSeek V4 Pro $235
Mistral Large 4 $356
Mistral Medium 3.5exceeds context $709
Claude Sonnet 5.5 $769
GPT-6.1 Sol $1,013
Gemini 3.1 Pro (preview) $1,071
Grok 4.7 $1,179
Claude Opus 5.5 $1,538
Claude Fable 5.1 $3,761
GPT-6 Astra $5,288

How the calculator counts

For each model: monthly cost = requests per day × 30 × (fresh input × input rate + cached input × cache rate + output × output rate), with the vendor's long-prompt multiplier applied to the whole request when the input crosses its threshold, and the batch multiplier applied when you tick batch pricing.

Sizes are converted between tokenizers using each vendor's own published rule of thumb. Words are converted at 5.7 characters per word including the space, which matches Google's "100 tokens is about 60-80 words".

TokenizerCharacters per tokenBasis
Claude 4.6 and earlier, Haiku 4.5 3.50 Anthropic glossary: a token is approximately 3.5 English characters. Source ↗
Claude 4.7 and later 2.69 Anthropic: the newer tokenizer produces approximately 30% more tokens for the same text. Source ↗
OpenAI 4.00 OpenAI's token-counting guide cites characters ÷ 4 as the common estimate for plain text. Source ↗
Gemini 4.00 Google: a token is about 4 characters; 100 tokens is about 60-80 English words. Source ↗
DeepSeek 3.33 DeepSeek: 1 English character is approximately 0.3 tokens. Source ↗
xAI 4.00 Not published by xAI; assumed equal to OpenAI's rule of thumb.
Mistral 4.00 Not published by Mistral; assumed equal to OpenAI's rule of thumb.

Not included: cache-write premiums (1.25× input on Anthropic and OpenAI, charged the first time a prefix is cached), cache storage fees on Google, priority and fast tiers, regional or data-residency surcharges (typically 10%), web-search and other tool fees, and image, audio or video input. DeepSeek is priced at off-peak rates with 21% of traffic assumed to fall in its weekday peak hours, which cost double; tick batch/off-peak pricing to price everything off-peak.

Rates come from the API price table and were checked against each vendor's pricing page on October 9, 2026. Models due to retire are flagged; see the retirement calendar for dates and replacements.

Frequently asked questions

How much does the Claude API cost per month?

It depends almost entirely on volume and caching. For a support chatbot handling 5,000 conversations a day (3,000 input and 400 output tokens each, 60% of input cached), Claude Sonnet 5.5 comes to about $1,467 a month. A coding agent running 800 turns a day over a 60K-token, 85%-cached context costs about $1,715 a month on Sonnet 5.5 and $3,431 on Opus 5.5. Both figures account for Claude's newer tokenizer, which counts roughly 30% more tokens than its older models for the same text.

Is GPT-6.1 Sol cheaper than Claude Sonnet 5.5?

Both list at $2 per million input tokens and $10 per million output, and since 7 October 2026 both charge $0.10 for cached input — Anthropic cut Sonnet 5.5's cache read from $0.20 that day, removing the one line on which Sol undercut it. What is left is the tokenizer: Claude's newer one counts about 30% more tokens for the same text. On the support-chatbot example above, GPT-6.1 Sol comes to $987 a month against $1,467 for Claude Sonnet 5.5. Sol's advantage shrinks or reverses on very long prompts, because OpenAI bills prompts over 272K tokens at twice the input rate and Sonnet 5.5 has no long-prompt surcharge at all.

What is the cheapest LLM API?

Of the current models in this calculator, GPT-6 Luna is cheapest for the support-chatbot workload at $51 a month, and Claude Fable 5.1 the most expensive at $7,233, a 143× spread for the same traffic. The cheapest model is rarely the right one: compare the price of the cheapest model that passes your own evaluations, and check whether a low rate is introductory (Gemini Flash rates double on 1 January 2027).

How can I reduce my LLM API bill?

Four levers move the number most. Cache the parts of the prompt that repeat (system prompts, tool definitions, documents), since cached input costs 90-95% less on most vendors. Send work that can wait through a batch API for 50% off at OpenAI, Anthropic and Google. Keep prompts under the long-context thresholds, where the rate jumps: 272K tokens at OpenAI and 200K at Google and xAI both double the input rate, and Claude Haiku 5.5 multiplies every rate by five above just 100K tokens. And route easy requests to a small model: the gap between a vendor's small and frontier tier is typically 10-100×.

Why does the calculator show different token counts for the same request?

Each vendor splits text into tokens differently, and you pay per token. Vendors publish rules of thumb: Anthropic says about 3.5 characters per token for Claude 4.6 and older and about 30% more tokens on Claude 4.7 and later, Google and OpenAI about 4 characters per token, and DeepSeek about 0.3 tokens per character. The calculator converts your sizes into each model's tokens using those figures. For a final decision, count a sample of your real prompts with each vendor's token-counting endpoint.