Claude Haiku 5.5 pricing
$0.10 input and $0.50 output per million tokens for prompts up to 100K tokens. Past 100K, the whole request moves to $0.50 / $2.50.
Verified 2026-10-09 against Anthropic's published rates · Compare with Luna 6 in the calculator
Rates
USD per million tokens
US-only inference (inference_geo: "us") multiplies every rate by 1.1. Prompt length includes cache reads and writes. Sources: Anthropic pricing, Claude Haiku 5.5 model overview.
What real workloads cost
See the full comparison of the pricing cliffs and workload costs inwhy Claude Haiku 5.5 can cost up to 12× more than GPT-6 Luna.
How the 100K threshold works
Anthropic prices Claude Haiku 5.5 by prompt length. Once a single request's prompt — new input plus cache reads plus cache writes — goes over 100,000 tokens, every token in that request, output included, bills at the higher rate: five times the base. A 100,001-token prompt costs about five times a 100,000-token one. Keeping retrieved context under 100K, or splitting long documents across requests, keeps you on the base tier.
How much does Claude Haiku 5.5 cost?
$0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. Above that, the whole request is billed at $0.50 input and $2.50 output.
Does the 100K threshold include cached tokens?
Yes. Anthropic counts every input token toward the prompt length, including cache reads and cache writes. A request with 60K cached and 50K new tokens is billed at the higher tier.
How much is prompt caching on Haiku 5.5?
Cache reads cost $0.01 per million tokens under 100K ($0.05 above). Cache writes cost $0.125 for a 5-minute cache or $0.20 for a 1-hour cache under 100K ($0.625 / $1.00 above).
Is there a batch discount?
Yes, the Message Batches API is 50% off: $0.05 / $0.25 under 100K and $0.25 / $1.25 above.
Compare before you commit
Run Haiku 5.5 and Luna 6 side by side in the playground.