Claude Haiku 3.5
$0.80 per million input tokens, $4 per million output. Superseded — kept here so you can price existing code.
| Input | $0.80 / MTok |
|---|---|
| Output | $4 / MTok |
| Cache read | $0.080 / MTok |
| Cache write (5 min) | $1 / MTok |
| Cache write (1 hour) | $1.60 / MTok |
| Batch API | $0.40 in / $2 out |
| Context window | 200,000 tokens |
| Tokenizer family | Claude (Sonnet 4.6 and earlier) |
| Tool-use overhead | 264 tokens added when any tool is defined |
What that costs per month
Monthly spend for a prompt of a given size, with a 500-token response, no caching. Rows are prompt size; columns are requests per day.
| Prompt size | 1,000 / day | 10,000 / day | 100,000 / day |
|---|---|---|---|
| 1,000 tokens | $85.17 | $851.67 | $8,517 |
| 5,000 tokens | $182.50 | $1,825 | $18,250 |
| 20,000 tokens | $547.50 | $5,475 | $54,750 |
Prompt caching on Claude Haiku 3.5
A 10,000-token static prefix at 10,000 requests a day costs $3,163 a month uncached. With caching at an 80% hit rate it costs $1,533 — a saving of $1,630, or 52%. Caching pays for itself after 1 read of the same prefix.
How cache economics workCheaper on input
Lower input price. Whether they are cheaper for your task depends on token counts and on whether they hold up on your evals.
- GPT-5.4 mini$0.75 / $4.50OpenAI
- Gemini 3 Flash Preview$0.50 / $3Google
- Gemini 3.5 Flash-Lite$0.30 / $2.50Google
- GPT-5 mini$0.25 / $2OpenAI
- Gemini 3.1 Flash-Lite$0.25 / $1.50Google
Price your actual prompt
These figures assume a prompt size. Paste your real one and the analyser will count it with Claude Haiku 3.5’s tokenizer family, find what is wasted, and compare against every other model at your volume.
Open the analyserPrices last verified 2026-08-10. Confirm against Anthropic’s own pricing page before committing to anything.