Claude API Pricing and Billing Explained (2026)

In this article

Claude API pricing is usage-based: you pay per token consumed across input, output, and optional cache reads. As of 2026, Claude Haiku 4.5 starts at $1 per million input tokens, Sonnet 5 at $2/M, and Opus 5 at $5/M. There are no seats or subscriptions on the API — only what you use. Rate limits scale with your usage tier, starting at Tier 1 ($5 spent) and rising through Tier 4 ($5,000 spent).

  • Pay-per-token: input tokens, output tokens, and optional prompt cache reads billed separately
  • Five-tier rate limit system based on cumulative spend (Tier 1–4 + Enterprise)
  • Prompt caching cuts repeated-context costs by up to 90% on cache reads

How much does Claude API cost per token?

Pricing varies by model family. Haiku is the cheapest and fastest; Opus is the most capable and most expensive. Output tokens cost more than input tokens across all models. Prompt cache reads are billed at a fraction of the standard input price, while cache writes carry a small surcharge to store the context. All figures below are from Anthropic's official pricing page.

ModelInput (per 1M tokens)Output (per 1M tokens)Cache ReadCache Write
Claude Haiku 4.5$1.00$5.00$0.10$1.25
Claude Sonnet 5$2.00$10.00$0.20$2.50
Claude Opus 5$5.00$25.00$0.50$6.25

Output tokens are 5x the price of input tokens on every current Claude model. If your workload generates long completions, that asymmetry dominates your bill. Batch API requests (asynchronous, up to 24-hour SLA) receive a 50% discount across all models — useful for offline evals or bulk document processing (Anthropic API docs).

What are the Claude API usage tiers?

Anthropic assigns new API accounts to Tier 1 by default and promotes accounts to higher tiers as cumulative spend crosses thresholds. Higher tiers unlock larger rate limits on both requests per minute (RPM) and tokens per minute (TPM). Promotion is automatic once you hit the threshold and hold a valid payment method for the required number of days.

TierSpend ThresholdWait PeriodClaude Sonnet 5 TPM
Tier 1$5 paid7 days after first charge40,000
Tier 2$50 paid7 days80,000
Tier 3$500 paid7 days160,000
Tier 4$5,000 paid14 days400,000
EnterpriseCustomNegotiatedCustom

Rate limits apply per-model and per-workspace. Hitting the TPM cap returns a 429 error. If you're seeing that frequently, you can either request a limit increase via the Anthropic Console or restructure requests to spread load — see our guide on handling Claude Code rate limit errors.

How does prompt caching reduce costs?

Prompt caching lets you mark a large, repeated prefix (system prompt, retrieved documents, conversation history) as cacheable. The first request writes the cache and costs 25% more than standard input pricing. Subsequent requests that hit the cache pay only 10% of the base input price. On workloads where you prepend the same 10,000-token system prompt on every call, cache reads cut that portion of costs by 90%.

Cache entries have a 5-minute TTL that resets on each hit. For agents making rapid follow-up calls (like Claude Code), the cache stays warm. For batch pipelines with long gaps between calls, factor in re-write costs. Full details in Anthropic's prompt caching docs.

How is API billing different from Claude Pro or Max?

Claude Pro ($20/month) and Max ($100–200/month) are subscription plans for claude.ai and Claude Code. They give you a soft usage cap — a rolling 5-hour token window — not a per-token meter. The API is purely pay-as-you-go with no monthly seat fee, but also no free usage beyond the initial trial credit.

  • API: per-token billing, no subscription, rate-limited by tier, no usage cap per se (just rate limits)
  • Claude Pro / Max: monthly subscription, soft 5-hour window cap, includes Claude Code access
  • Claude Code with API key (BYO key): routes Claude Code through your own API key, billed as API usage at standard token rates

Developers using Claude Code with their own Anthropic API key are billed at API rates, not subscription rates. A single long agentic session can consume millions of tokens, so the meter runs faster than it looks. See our post on why Claude Code uses so many tokens and how to reduce Claude Code token usage.

How do you read your Anthropic API bill?

Billing is itemized in the Anthropic Console under Usage. Each line shows model, date, input tokens, output tokens, cache read tokens, and cache write tokens with their respective per-unit costs. There is no invoice PDF until you download it — the dashboard is the primary source of truth. Usage data updates with roughly a 1-hour lag.

For teams running Claude Code under a Max subscription, usage tracking works differently: you watch the 5-hour rolling window and the weekly cap, not a dollar meter. The Usagebar menu bar app ($9 one-time) shows your current 5-hour window burn rate, weekly cap, and context size at a glance, with notifications at 50%, 75%, and 90% of your limit — before you hit a hard stop mid-session. You can also use the /usage slash command inside Claude Code or check claude.ai/settings/usage.

What are the best ways to lower Claude API costs?

  • Use Haiku for classification and routing tasks. Haiku costs 20% of Opus on both input and output. Triaging requests before escalating to Sonnet or Opus cuts costs significantly.
  • Enable prompt caching on any repeated context block over ~1,000 tokens. Cache reads are 90% cheaper than fresh input tokens.
  • Use Batch API for async workloads. 50% discount in exchange for up to 24-hour turnaround.
  • Shorten system prompts. Every token in a system prompt is re-billed on each call if caching is off.
  • Set max_tokens carefully. Over-provisioning does not cost money if the model stops early, but it can waste compute on verbose completions. Use concise output instructions.
  • Monitor context window growth in agents. Agentic loops accumulate conversation history. Truncating or summarizing older turns periodically keeps per-call costs flat. See our post on how to check Claude Code token count.

What is the Claude API context window limit?

As of 2026, Claude Opus 5 and Sonnet 5 support a 1,000,000-token context window, while Haiku 4.5 supports 200,000 tokens. You are billed for every token in the context on each API call — context is not free just because it was sent in a prior turn. For long-running agent sessions, context accumulation is the primary cost driver, not per-response output length.

Key takeaways for Claude API pricing

  1. Input, output, cache read, and cache write tokens are billed separately — the rates differ significantly.
  2. Output tokens cost 5x input tokens; optimize for short completions where possible.
  3. Prompt caching saves up to 90% on repeated context — worth enabling for any system prompt over 1,000 tokens.
  4. Rate limits depend on usage tier (Tier 1–4), not plan type. Tiers advance automatically as spend accumulates.
  5. Batch API gives a 50% discount for async workloads with 24-hour SLA tolerance.
  6. API billing and subscription billing (Pro/Max) are separate products — using Claude Code with a BYO API key runs at token rates, not subscription rates.
  7. If you're on a Max subscription using Claude Code, track your 5-hour window and weekly cap with the /usage command or Usagebar to avoid mid-session cutoffs.

Sources

Never Get Locked Out Mid-Task Again

Never hit your usage limits unexpectedly. Usagebar lives in your menu bar and shows your 5-hour and weekly limits at a glance.

Get Usagebar

$9 — one-time, lifetime updates