Skip to content
Antonio Kodheli

Claude API Pricing (October 2026): Every Model, With Real Cost Examples

Claude API pricing for every current model: Fable 5.1, Opus 5.5, Sonnet 5.5 and Haiku 4.5 per million tokens, plus prompt caching, Batch API discounts, tool costs and worked examples.

Antonio Kodheli··7 min read

Claude API pricing is per token, and it depends on the model. As of October 6, 2026, the cheapest current model, Claude Haiku 4.5, costs $1 per million input tokens and $5 per million output tokens. Claude Sonnet 5.5 costs $2 / $10, Claude Opus 5.5 costs $4 / $20, and Claude Fable 5.1, the most capable model, costs $10 / $50. Prompt caching and the Batch API can cut those numbers by half or more.

All prices below come from Anthropic's pricing page, checked on October 6, 2026. Prices change, so check that page before you commit a budget. If you are just starting, here is how to get a Claude API key.

Claude API pricing by model

Prices are in US dollars per million tokens (MTok).

ModelAPI model IDInputOutputCache hit
Claude Fable 5.1claude-fable-5-1$10$50$0.25
Claude Opus 5.5claude-opus-5-5$4$20$0.20
Claude Sonnet 5.5claude-sonnet-5-5$2$10$0.20
Claude Haiku 4.5claude-haiku-4-5-20251001$1$5$0.10

Earlier models are still on the API at their own rates:

ModelInputOutput
Claude Fable 5$10$50
Claude Opus 5, 4.8, 4.7, 4.6 and 4.5$5$25
Claude Sonnet 5$2$10
Claude Sonnet 4.6 and 4.5$3$15

A few notes that matter when you compare:

  • Output costs five times input on every current model. Long answers cost more than long prompts.
  • Sonnet 5's $2 / $10 price is now permanent. It launched as introductory pricing, and the planned rise to $3 / $15 in September 2026 was cancelled.
  • Newer models count more tokens for the same text. Claude 4.7 and later use a newer tokenizer that produces roughly 30% more tokens for the same text. Compare models on cost per task, not just the per-token price.
  • Claude Mythos models are priced like Fable but offered by invitation only.

What is a token?

A token is a chunk of text the model reads or writes. Anthropic's rule of thumb is that 1 token is about 4 characters, or 0.75 English words, so 1 million tokens is roughly 750,000 words. Everything you send counts as input: the system prompt, the conversation history, documents and tool definitions. Everything the model writes counts as output.

Prompt caching: the biggest saving

If you send the same long prefix on every request (a system prompt, a document, a tool list), prompt caching stores it so later requests read it cheaply:

Cache operationPrice
5-minute cache write1.25 × base input
1-hour cache write2 × base input
Cache hit (read)0.1 × base input (0.05 × on Opus 5.5, 0.025 × on Fable 5.1)

A 5-minute cache pays for itself after one read, and a 1-hour cache after two. For a chatbot with a long system prompt, caching often cuts the input bill by more than half.

Batch API: 50% off anything that can wait

The Batch API processes requests asynchronously at half price on both input and output. Batch rates for current models:

ModelBatch inputBatch output
Claude Fable 5.1$5$25
Claude Opus 5.5$2$10
Claude Sonnet 5.5$1$5
Claude Haiku 4.5$0.50$2.50

Use it for anything nobody is waiting on: overnight summaries, classification, data extraction, evaluations. Batch and caching discounts stack.

Other costs to budget for

  • Web search: $10 per 1,000 searches, plus the tokens for the results.
  • Web fetch: no extra charge beyond the tokens of the fetched page.
  • Code execution: free when used with web search or web fetch. Otherwise each organization gets 1,550 free hours a month, then $0.05 per hour per container.
  • Tool definitions: using tools adds a small system prompt (about 290 tokens on Opus 5.5 and Sonnet 5.5) plus the size of your tool schemas, billed as input.
  • US-only inference: pinning inference to the US with inference_geo: "us" multiplies all token prices by 1.1 on Claude 4.6 and later.
  • Fast mode (research preview, Opus only): Opus 5.5 at $8 / $40 per million tokens for faster output.
  • Long context: Claude 4.6 and later models get the full 1M-token context window at standard per-token rates.

Claude API cost examples

These are my own calculations from the rates above.

1. A support chatbot on Sonnet 5.5. Each reply sends about 3,000 input tokens (system prompt, history and the question) and gets about 400 back. 3,000 × $2/M + 400 × $10/M = $0.010 per reply, or about $100 per 10,000 replies. If 2,500 of those input tokens are a cached system prompt: 500 × $2/M + 2,500 × $0.20/M + 400 × $10/M = $0.0055 per reply, about $55 per 10,000 (plus a small cost for the occasional cache write).

2. The same chatbot on Haiku 4.5. 3,000 × $1/M + 400 × $5/M = $0.005 per reply, about $50 per 10,000 before caching. For simple, well-scoped questions, Haiku is often good enough.

3. Summarising 1,000 long documents with Opus 5.5. 20,000 tokens in and 1,000 out per document is 20M input and 1M output tokens. Standard: 20 × $4 + 1 × $20 = $100. Batch API: 20 × $2 + 1 × $10 = $50.

4. One hard reasoning task on Fable 5.1. 50,000 tokens in and 5,000 out: 50,000 × $10/M + 5,000 × $50/M = $0.75 per request. Worth it for the problems only the strongest model solves; expensive as a default.

How to keep Claude API costs down

  1. Pick the smallest model that passes your tests. Anthropic's own guidance is Haiku for simple tasks, Sonnet for most production work and Opus for the hardest reasoning.
  2. Cache everything that repeats. System prompts, documents and tool lists.
  3. Batch what can wait. Half price for a few hours of patience.
  4. Cap the output. Set max_tokens, and ask for concise formats; output is the expensive side.
  5. Trim the context. Send the relevant part of a document, not the whole thing, and summarise long conversation history.
  6. Measure per task. Log input, output and cache tokens per feature so you know what each one costs.

Claude API pricing vs a Claude subscription

A Claude Pro or Max subscription covers using Claude in the apps (and Claude Code). The API is billed separately, pay-as-you-go, in the Claude Console. New API accounts get a small amount of free credit to test with; after that you pay per token. If you are building agents, the Claude Agent SDK bills through the same API account.

Frequently asked questions

How much does the Claude API cost? From $1 per million input tokens and $5 per million output tokens on Haiku 4.5, up to $10 / $50 on Fable 5.1. Most production apps run on Sonnet 5.5 at $2 / $10.

Is the Claude API free? No, but new accounts get a small free credit to test with. After that it is pay-as-you-go per token.

Which Claude model is cheapest? Claude Haiku 4.5, at $1 / $5 per million tokens, or $0.50 / $2.50 through the Batch API.

Is the Claude API cheaper than a subscription? It depends on volume. Light personal use is usually cheaper on a subscription; an app serving many users needs the API, and caching and batching keep it affordable.

Do prompt caching and the Batch API stack? Yes. Their multipliers combine, along with any data residency multiplier.


Want a Claude feature in your product without a surprise bill? I design for cost from the start: model choice, caching, batching and per-feature tracking. See my AI integration services.

I build Claude features into products and design them to stay affordable at scale: model choice, caching and batching from day one.

Get an AI feature priced

Written by Antonio Kodheli

Full-stack web developer in Boston. I build modern web apps with TypeScript, Next.js, and the Claude API, and write about it here.