Anthropic's new small model costs a tenth of Haiku 4.5 per token, matches GPT-6 Luna's price, and has one pricing cliff worth knowing about
Haiku 5.5 costs $0.10 per million input tokens. Anthropic launched Claude Haiku 5.5 on October 7, 2026 at that price plus $0.50 per million output tokens, for prompts up to 100,000 tokens - Anthropic. That is 90% below Haiku 4.5's list price and exactly the price of GPT-6 Luna, the small model OpenAI shipped in late September. Anthropic says the new model costs around 75% less to run than its predecessor once everything is counted, and its launch table puts it ahead of Luna on every test the two share.
The catch is in the fine print. Haiku 5.5 is the only model on Anthropic's current price list that is priced by prompt length, its new tokenizer counts the same text as more tokens, and the independent numbers tell a more mixed story than the launch table. This dispatch covers what shipped, where the 75% figure comes from, how Haiku 5.5 stacks up against GPT-6 Luna, and what to change in your app this week.
1. What Anthropic Shipped on October 7
The headline is price, but the model itself also grew. Haiku 5.5 has a 1M token context window and 128K max output, up from 200K and 64K on Haiku 4.5, and it is the first Haiku with adaptive thinking and the effort parameter - Claude Docs. It is live as claude-haiku-5-5 on the Claude API, Amazon Bedrock, Google Cloud and Microsoft Foundry, with retirement promised no sooner than October 7, 2027 - Claude Docs.
Three other changes landed the same day. Sonnet 5.5 cache reads dropped from $0.20 to $0.10 per million tokens - Claude release notes. Max and Team subscribers began receiving a monthly API credit of $100 (Max 5x), $200 (Max 20x), or $20 to $100 per Team seat pooled up to $500 - Claude Help Center. And the Python and TypeScript SDKs gained beta browser use and computer use toolset classes you subclass against your own browser or desktop - Claude Docs.
| Model | Input / 1M | Output / 1M | Cache read / 1M |
|---|---|---|---|
| Claude Haiku 5.5 (prompt up to 100K) | $0.10 | $0.50 | $0.01 |
| Claude Haiku 5.5 (prompt over 100K) | $0.50 | $2.50 | $0.05 |
| Claude Haiku 4.5 | $1.00 | $5.00 | $0.10 |
| GPT-6 Luna (prompt up to 272K) | $0.10 | $0.50 | $0.01 |
| Claude Sonnet 5.5 | $2.00 | $10.00 | $0.10 |
The Claude rows are Anthropic's list prices - Anthropic pricing. The Luna row is OpenAI's list price - OpenAI. Why this matters: for the first time, the two biggest labs sell their small models at an identical base price, so the choice between them comes down to quality, speed and the edges of the rate card rather than the headline number.
2. Where the 75% Comes From
The 75% is a blend of two very different numbers. Anthropic prices Haiku 5.5 90% below Haiku 4.5 on prompts up to 100,000 tokens and 50% below on longer ones, and notes that about 90% of Haiku 4.5 requests fell under that threshold - Anthropic. But the model uses the newer Claude tokenizer, so the same text counts as roughly 30% more tokens than on Haiku 4.5 - Claude Docs.
Do the arithmetic from first principles and the picture sharpens. A short request now costs about 0.10 x 1.3, or 13% of what it did, an 87% saving. A long request costs 0.50 x 1.3, or 65% of before, only a 35% saving. The tokenizer also moves the threshold itself: a prompt that measured around 80,000 tokens on Haiku 4.5 can count past 100,000 on Haiku 5.5 and jump to the higher tier. Long requests carry far more tokens than short ones, which is how a 90% list cut on nine requests in ten can average out near 75%.
The comparison with Luna follows the same logic. OpenAI's long-context surcharge starts at 272K input tokens and only doubles input to $0.20 - OpenAI. So between 100K and 272K tokens, Haiku 5.5 costs five times as much per token as Luna. The chart prices a request with 2,000 output tokens at each prompt size, at the same token count on both models (real counts depend on each model's tokenizer).
How to apply this: if your Haiku workload is classification, routing or short summaries, the saving is real and close to the full 87%. If you feed it long documents, measure your prompts on the new tokenizer first, and lean on prompt caching, since Haiku 5.5 cache reads cost $0.01 per million below the threshold. We broke that technique down in our guide to cutting AI costs with prompt caching.
3. Haiku 5.5 vs GPT-6 Luna on Benchmarks
Anthropic's launch table compares Haiku 5.5 directly with GPT-6 Luna, and the gaps are wide. On OSWorld 2.1 (computer use) Haiku 5.5 scores 72.4% against 48.9%, on Terminal-Bench 4.0 (agentic coding) 39.2% against 16.4%, and on the GDPval-AA knowledge-work Elo 1620 against 1437 - Anthropic. These are vendor-run results, so treat them as a ceiling rather than a verdict.
The independent numbers agree on direction but add a cost. Artificial Analysis scores Haiku 5.5 at max effort 43 on its Intelligence Index v4.3.2, at 244 output tokens per second and $0.21 per index task - Artificial Analysis. GPT-6 Luna at max effort scores 38, at 127 tokens per second, for $0.07 per task - Artificial Analysis. The per-token price is the same, so at max effort Haiku uses about three times the tokens per task. Smarter and faster, yes. Cheaper per finished task at the top setting, no.
That is why the effort setting matters more than the rate card. Anthropic defaults Haiku 5.5 to medium and recommends low for chat, short tool tasks and high-volume requests - Claude Docs. We covered how to tune that dial in our guide to setting the effort dial to cut AI costs. For a hands-on look, Bijan Bowen's 31-minute test puts Haiku 5.5 through a browser OS build and a series of game builds, with a usage overview at 28:10.
4. What to Change in Your App This Week
Haiku 5.5 is not a drop-in swap for Haiku 4.5. Code written for the old model can return errors on the new one, and Anthropic's migration guide lists each breaking change with a before-and-after request - Claude Docs. Most teams will hit at least one of them on the first request, because the defaults changed as well as the limits: adaptive thinking is now on unless you turn it off.
Every fix is small and mechanical, which is good news for a model you will probably route a lot of traffic to. The two that return a 400 error outright, and so break a working integration the moment you change the model ID, are worth fixing before anything else:
- Remove
budget_tokensand sampling parameters: manual extended thinking and any non-defaulttemperature,top_portop_kare rejected - End on a user turn: assistant prefill is no longer accepted
The quieter changes cost money rather than uptime. Responses can now open with thinking blocks, so code that reads the first content block as the answer needs to select blocks by type. Thinking tokens also count toward max_tokens, so a small limit can stop after the thinking block and before any text, and every token budget shifts with the new tokenizer. Set the limit with room for thinking, recount your prompts, then pick the lowest effort level your evals accept.
If your app routes cheap work to a small model and hard work to a larger one, Haiku 5.5 moves where that line sits, the trade-off we mapped in our guide to cutting agent costs with model routing. For the parallel changes on the bigger models, see our Claude 5.5 migration guide.
The new API credit covers the Claude API, Message Batches, Managed Agents and the Agent SDK, but not interactive Claude Code and not Claude on Bedrock, Vertex AI or Foundry, and unused credit expires each cycle - Claude Help Center. At $0.10 per million input tokens, a $100 credit buys about a billion short-prompt input tokens on Haiku 5.5.
On Founden, the change already landed. Founden's model catalog, which refreshes from vendor feeds, listed Haiku 5.5 at 23:14 UTC on launch day, and because Founden's model choices follow model families rather than pinned versions, a builder set to Claude's Haiku family now resolves to Haiku 5.5 without anyone touching a setting.
5. The Bottom Line
Haiku 5.5 resets what a small model costs and what it can do. On Anthropic's tests and on Artificial Analysis's index, it is the stronger and faster model at Luna's price, and by our arithmetic short prompts run about 87% cheaper than Haiku 4.5. The structural reason is simple: both labs now sell small-model tokens at the same price, so the competition has moved to intelligence per token and to the edges of the rate card.
Use that to decide. Switch to Haiku 5.5 for high-volume short work (classification, extraction, routing, subagents), where the price is identical and Anthropic's widest leads are in computer use and agentic coding. Benchmark GPT-6 Luna for prompts between 100K and 272K tokens, where Luna is five times cheaper per token, and for max-effort batch jobs where cost per finished task matters more than raw score. Whichever you choose, set effort deliberately, because the dial now moves your bill more than the model name does. For how small-model costs feed into product pricing, see our guide to pricing your AI product to beat token costs.
This dispatch reflects pricing and benchmarks as of October 8, 2026. Model prices and limits change often, so check the vendors' pricing pages before you budget.