Haiku 5.5 costs a tenth of Haiku 4.5. Swap only the model ID and your classifier loses the tickets it most needs to get right
Claude Haiku 5.5 shipped on 7 October at $0.10 in and $0.50 out per million tokens, against $1 and $5 for Haiku 4.5. We ran a five-label ticket classifier written for Haiku 4.5 through it 45 times per setting. With only the model ID changed, 7 of 45 replies had no usable label, and the misses were the ambiguous tickets, every time. One request field brought it back to 45 of 45 at $9.50 per million tickets. Here are the measurements, the two requests that now return 400, and the 100,000-token price threshold that is counted in the new tokenizer.
Claude Haiku 5.5 went live on 7 October with a list price of $0.10 per million input tokens and $0.50 per million output tokens, for prompts up to 100,000 tokens. Haiku 4.5 costs $1 and $5. For anyone running classification, routing or extraction at volume, that is the reason to move this week, and Claude Code 2.1.293 already made it the default Haiku, and GitHub Copilot added it the same day.
The migration guide lists ten changes. On 8 October we took a classifier written the way most Haiku 4.5 classifiers are written and measured which of those changes bite, and where.
The classifier we tested
SYSTEM = ("Classify the customer message into exactly one label: "
"billing, bug, feature_request, account, other. Reply with the label only.")
resp = client.messages.create(
model="claude-haiku-4-5",
max_tokens=10,
temperature=0,
system=SYSTEM,
messages=[{"role": "user", "content": ticket}],
)
label = resp.content[0].text.strip()
Fifteen support tickets: five one-liners, a handful of messy multi-issue ones, two security reports with an injection payload in the text, and one about a leaked API key. Each configuration ran all fifteen three times, so 45 requests per row. A reply counts as a label only if the text is exactly one of the five strings. Every request in this post together cost under ten cents.
Two parts of that snippet fail before anything is classified. temperature=0 returns a 400 with the message `temperature` is deprecated for this model. over raw HTTP, and on the Python SDK 1.12.1 it raises TypeError: Messages.create() got an unexpected keyword argument 'temperature' without sending a request. The temperature post covers why that parameter is gone across vendors. And if your extractor ends messages with an assistant turn such as "Label:" or "{", Haiku 5.5 answers This model does not support assistant message prefill. The conversation must end with a user message. Haiku 4.5 accepted both.
Delete temperature and the request goes through. That is where the quiet part starts.
What came back with only the model ID changed
| Setting | Valid label | Empty | Prose | Avg input tokens | Avg output tokens | Cost per 1M tickets | Median latency |
|---|---|---|---|---|---|---|---|
Haiku 4.5, max_tokens 10 | 45 | 0 | 0 | 56.6 | 4.3 | $77.93 | 0.53 s |
Haiku 5.5, ID swap, max_tokens 10 | 38 | 5 | 2 | 76.3 | 4.7 | $9.98 | 0.69 s |
Haiku 5.5, max_tokens 1024 | 43 | 0 | 2 | 76.3 | 16.0 | $15.66 | 0.69 s |
Haiku 5.5, effort low, max_tokens 10 | 43 | 0 | 2 | 76.3 | 4.1 | $9.67 | 0.63 s |
Haiku 5.5, thinking disabled, max_tokens 10 | 45 | 0 | 0 | 75.3 | 3.9 | $9.50 | 0.73 s |
Haiku 5.5, JSON schema + effort low, max_tokens 32 | 45 | 0 | 0 | 301.3 | 10.9 | $35.57 | 1.31 s |
Cost is list price times measured tokens, with no caching or batch discount. Each row is one run of 45; repeating a row moved individual counts by one or two, and the ID-swap row never reached 45.
The missing labels on the second row are not random. Across every run we logged, they came from three tickets, all of them about two things at once: an SSO migration that stripped admin rights while the invoice kept billing for them, the leaked-API-key question, and once a GDPR deletion request that also asked for a refund. On those, Haiku 5.5 decided to think. Adaptive thinking is on by default at medium effort, thinking tokens count toward max_tokens, and ten tokens ran out inside the thinking block. The response has stop_reason: "max_tokens" and a single content block of type thinking, whose thinking field is an empty string because Haiku 5.5 omits the thinking text by default. On the Python SDK, the last line of the classifier then raises:
AttributeError: 'ThinkingBlock' object has no attribute 'text'
The five one-line tickets never thought in any run. So a migration test on a random sample of easy traffic passes, and the failures are concentrated on exactly the inputs a human would also hesitate over. The SSO ticket failed five out of five times at the default setting.
Raising max_tokens to 1024 stops the crash and exposes a second problem. Two replies were prose, both on the leaked-key ticket. In a separate run of the same request, one read, in full: "Which label fits best: billing, bug, feature_request, account, or other?\n\nHmm, wait. I should give just the label, so let me answer directly: other". Another ended on account for the same ticket after two paragraphs of reasoning. Anthropic's prompting guide for Haiku 5.5 names this behaviour, reasoning-like text in the visible reply, and says it appears more often with thinking off or at low effort. In our runs it went the other way: thinking off produced none, low and the default each produced two. Treat that as one prompt on one day; the general lesson is that the reply format is no longer guaranteed by the instruction alone.
The thinking also sits where the cost estimate does not look. Across the 45 requests at 1024, average output was 16 tokens, but a ticket that thought produced between 104 and 214 output tokens in our runs. At $0.50 per million, a ticket that thinks for 160 tokens costs $80 per million such tickets in output alone, more than the roughly $78 per million that Haiku 4.5 charged for the whole request.
The setting that matches what Haiku 4.5 was doing
Haiku 4.5 ran a classifier without thinking unless you asked for it. The closest Haiku 5.5 equivalent is to say so:
resp = client.messages.create(
model="claude-haiku-5-5",
max_tokens=10,
thinking={"type": "disabled"},
system=SYSTEM,
messages=[{"role": "user", "content": ticket}],
)
text = "".join(b.text for b in resp.content if b.type == "text").strip()
if text not in LABELS:
route_to_fallback(ticket, resp)
That row is 45 of 45 at $9.50 per million tickets, about an eighth of the Haiku 4.5 cost on the same traffic. thinking: {"type": "disabled"} is accepted at low, medium and high effort and returns a 400 at xhigh and max. The Models API reports per model whether it is accepted, in capabilities.thinking.types.disabled, added on 5 October.
Lowering effort to low is the option the migration guide points to first, and on this workload it was not enough: no empty replies, but two prose answers on the hardest tickets. The guide also says telling the model in the prompt to answer directly did not stop it from thinking in Anthropic's own tests, so "Reply with the label only" in the system prompt is not a substitute for the request field.
A JSON schema with an enum of the five labels also reached 45 of 45 at low effort, and it is the right tool when the output has several fields. For a single label it is the expensive option here: structured output added about 225 input tokens per request and doubled the median latency, which took the cost to $35.57 per million tickets. At the default effort the same schema with max_tokens 32 returned an empty reply on 8 of 10 tries on the two hard tickets, so a schema does not remove the need to leave room for thinking or turn it off.
Whatever you pick, keep the last two lines of that snippet. Select text blocks by type, never by position, and check the result against the label set before it reaches anything downstream. That check is what turns a silent misroute into a counted fallback.
The token count moved, and so did a price threshold
Haiku 5.5 uses the tokenizer introduced with Claude Opus 4.7. Anthropic gives the increase as roughly 30% for the same text; our system prompt plus ticket went from 55 tokens on Haiku 4.5 to 73 on Haiku 5.5, and the 15-ticket average from 56.6 to 76.3, about 35%. Anything budgeted in tokens changes with it: max_tokens limits, chunk sizes for extraction, and cost dashboards that divide by a token count measured on the old model. Recount with /v1/messages/count_tokens and "model": "claude-haiku-5-5" rather than reusing old numbers.
The token count matters twice for Haiku 5.5, because it is the one current Claude model priced by prompt length. Up to 100,000 input tokens it costs $0.10 in, $0.50 out, $0.01 for a cache read. Above 100,000 every line is five times higher: $0.50, $2.50 and $0.05. The threshold is measured in Haiku 5.5 tokens. A long-document extraction prompt that counted 80,000 tokens on Haiku 4.5 is likely to count about 104,000 on Haiku 5.5, and it is billed at the higher tier. That is still half of what Haiku 4.5 charges, so nobody pays more than before; a cost model built on the $0.10 headline is off by a factor of five on those requests. If your documents cluster near the line, chunk to stay under it, and check the per-request usage.input_tokens rather than an average.
The rest of the list, in the order it is likely to matter
- Refusals are new on Haiku. Haiku 5.5 runs safety classifiers that can return
stop_reason: "refusal"with astop_details.categoryofcyber,frontier_llm,bioorgeneral_harms, and there is no server-side fallback. Our two security tickets with SQL injection and XSS payloads were labelledbugwith no refusal, but a triage pipeline that receives exploit code from users should handle the stop reason anyway, because a retry of the same request usually refuses again. - Priority Tier does not cover Haiku 5.5. An organisation with a Priority Tier commitment on Haiku 4.5 has to plan that capacity separately.
- Forced tool choice still works on Haiku.
tool_choiceofanyor a named tool is accepted on Haiku 5.5, unlike Sonnet 5.5 and Fable 5.1, but the response then starts with the tool call and contains no thinking. - Conversations that edit history. If you send thinking blocks back and also change
system,toolsor earlier messages between turns, the request returns a 400. Keep those conversations append-only. - Computer use moved to a toolset.
computer_20250124returns a 400 on the Claude API and Google Cloud; the replacement iscomputer_toolset_20260801. - The ID has no alias.
claude-haiku-5-5is a fixed ID with no date suffix, so there is nothing to pin beyond it.
There is no deadline. Haiku 4.5 is not deprecated; the deprecations page lists it as active with retirement "not sooner than October 15, 2026", and Anthropic commits to at least 60 days of notice. The move is worth making for the price, and it is safe to make once the ambiguous tail of your own traffic passes. Pull the fifty tickets your team argued about last month, run them through the new request a few times each, and count valid labels. A random sample of easy traffic will tell you the migration worked when it has not.
For the vocabulary, see tokenizer, token cost, reasoning model and eval. For the model itself, Claude and the ChatGPT vs Claude comparison.
Get the next post when it ships
One email on Sunday with the new post and a short list of what shipped that week — new guides, tool updates, and a couple of links worth reading.