News Claude
Anthropic launches Claude Haiku 5.5: benchmarks and new pricing

Claude Haiku 5.5 is available now. Anthropic released it on October 7 for high-volume work: extracting data, classifying requests, summarizing documents and handling scoped tasks within agents. API pricing starts at $0.10 per million input tokens and $0.50 per million output tokens, for prompts up to 100,000 tokens.
The official announcement also brings changes for people using other Claude models: Sonnet 5.5 cache reads at half the price, and new monthly API credits for Max and Team. This story’s cover image is from Anthropic.
Haiku 5.5 vs. Haiku 4.5, GPT-6 Luna and Sonnet 5.5
Choose a test to compare the launch results. GDPval-AA and AA-Briefcase use Elo ratings; the other tests show percentages. Each measures different tasks: compare models within the same benchmark.
Compare by task
GDPval-AA v2.1
Higher is betterProfessional deliverables across 44 occupations. Elo compares results; it is not an accuracy percentage.
- Haiku 5.5
- 1620
- Haiku 4.5
- 735
- GPT-6 Luna
- 1437
- Sonnet 5.5Reference
- 1840
Independently evaluated by Artificial Analysis. Haiku 5.5 scores 1620 at max effort; at medium, the default, it scores 1277.
AA-Briefcase v1.1
Higher is betterLong-running professional projects with linked tasks and many source files. Elo score.
- Haiku 5.5
- 1578
- Haiku 4.5
- 614
- GPT-6 Luna
- 1336
- Sonnet 5.5Reference
- 1824
Independently evaluated by Artificial Analysis. Haiku 5.5 scores 1372 at medium effort and 1578 at max.
OSWorld 2.1
Higher is betterComputer use, offline subset. These are not OSWorld 2.0 results or results for the full task set.
- Haiku 5.5
- 72.4%
- Haiku 4.5
- 15.7%
- GPT-6 Luna
- 48.9%
- Sonnet 5.5Reference
- 83.9%
Anthropic evaluated GPT-6 Luna on the same 82 tasks. The system card distinguishes partial-credit scoring from passing every checkpoint in a task.
Humanity’s Last Exam · no tools
Higher is betterMultidisciplinary academic knowledge and reasoning, without tools.
- Haiku 5.5
- 45.9%
- Haiku 4.5
- 10.2%
- GPT-6 Luna
- —
- Sonnet 5.5Reference
- 56.9%
The table does not publish a GPT-6 Luna result for this test. A dash means no reported result.
Humanity’s Last Exam · with tools
Higher is betterThe same evaluation family with tools available. This is a different configuration from the previous chart.
- Haiku 5.5
- 57.4%
- Haiku 4.5
- 18.7%
- GPT-6 Luna
- —
- Sonnet 5.5Reference
- 64.5%
Do not combine results with and without tools into one score. There is no published GPT-6 Luna result here.
Terminal-Bench 4.0
Higher is betterAgentic coding tasks in a terminal.
- Haiku 5.5
- 39.2%
- Haiku 4.5
- 0.0%
- GPT-6 Luna
- 16.4%
- Sonnet 5.5Reference
- 70.6%
Haiku 4.5's 0.0% is the reported result. Haiku 5.5 improves, but Sonnet 5.5 retains a substantial lead on this test.
FrontierCode 1.1 · Main
Higher is betterThe 100 hardest tasks in Cognition's coding evaluation.
- Haiku 5.5
- 46.4%
- Haiku 4.5
- —
- GPT-6 Luna
- 42.4%
- Sonnet 5.5xhigh
- 52.1%
The launch table uses Sonnet 5.5 at xhigh: 52.1%. The system card also reports 46.2% at max. Haiku 5.5 uses max; no Haiku 4.5 result is listed.
Chartography
Higher is betterVisual reasoning about charts, without tools.
- Haiku 5.5
- 46.4%
- Haiku 4.5
- 6.4%
- GPT-6 Luna
- 29.1%
- Sonnet 5.5Reference
- 61.6%
This evaluation does not measure copy quality or campaign performance.
View the full table
| Benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| GDPval-AA v2.1 (Elo) | 1620 | 735 | 1437 | 1840 |
| AA-Briefcase v1.1 (Elo) | 1578 | 614 | 1336 | 1824 |
| OSWorld 2.1 (%) | 72.4% | 15.7% | 48.9% | 83.9% |
| Humanity’s Last Exam · no tools (%) | 45.9% | 10.2% | — | 56.9% |
| Humanity’s Last Exam · with tools (%) | 57.4% | 18.7% | — | 64.5% |
| Terminal-Bench 4.0 (%) | 39.2% | 0.0% | 16.4% | 70.6% |
| FrontierCode 1.1 · Main (%) | 46.4% | — | 42.4% | 52.1% (xhigh) |
| Chartography (%) | 46.4% | 6.4% | 29.1% | 61.6% |
The improvement over Haiku 4.5 spans very different tasks. On OSWorld 2.1 offline, its score rises from 15.7% to 72.4%. On Terminal-Bench 4.0 it reaches 39.2%, against the 16.4% reported for GPT-6 Luna; Sonnet 5.5 scores 70.6%. That gap matters when choosing an agent for complex coding work.
The configuration changes the comparison
The system card, section 8, explains the environments and reasoning effort. Haiku 5.5 is generally evaluated at max, while the API defaults to medium. On GDPval-AA it scores 1620 Elo at max and 1277 at medium; at medium it uses about a tenth of the output tokens. A low token price and a benchmark score alone therefore cannot tell you what your task will cost.
There is another visible difference in FrontierCode: the announcement lists Sonnet at 52.1% with xhigh. The system card also reports 46.2% at max. Our chart retains the launch value and its setting. Dashes mean no reported result; Haiku 4.5’s 0.0% on Terminal-Bench is an actual published value.
Haiku 5.5 API pricing: two prompt-length tiers
The pricing documentation sets the tier by prompt length. Above 100,000 tokens, the request’s input, output and caching rates change.
| Usage | Haiku 5.5 | Haiku 4.5 | Sonnet 5.5 | |
|---|---|---|---|---|
| Up to 100,000prompt tokens | Over 100,000prompt tokens | |||
| Input | $0.10 | $0.50 | $1.00 | $2.00 |
| Output | $0.50 | $2.50 | $5.00 | $10.00 |
| Cache read | $0.01 | $0.05 | $0.10 | $0.10 |
| Cache write · 5 min | $0.125 | $0.625 | $1.25 | $2.50 |
| Cache write · 1 h | $0.20 | $1.00 | $2.00 | $4.00 |
Scroll across to compare all four rates.
Per-token pricing is 90% lower than Haiku 4.5 in the short-prompt tier and 50% lower in the long tier. Anthropic estimates about 75% lower cost per task on average, accounting for its request mix and changes in token consumption. That average does not promise identical savings for every integration.
One message, two types of tokens
You write a message and Claude replies. The API counts each text separately: input for your message, output for the reply. A token is a piece of text; a word can take more than one.
Write a short reply thanking a customer for their comment.
Thanks for sharing your experience! We're glad it was helpful.
Cost per request
Prompt up to 100,000 tokens
Enter whole token counts within the model limits.
- Haiku 5.5
- $0.000028$0.028 per 1,000 identical requests
- Haiku 4.5
- $0.00028$0.280 per 1,000 identical requests
- Sonnet 5.5
- $0.00056$0.560 per 1,000 identical requests
The example uses 80 input tokens and 40 output tokens. Text and counts are illustrative; each model may count and respond differently. Prices use API rates; they do not calculate your Claude subscription bill.
Adjust the token counts
Standard API rates in USD, excluding caching and tool charges. Thinking tokens also count as output. Compares token volumes; the cost of a task depends on what each model consumes.
Official rates ↗The calculator lets you try the jump from 100,000 to 100,001 tokens. It compares equal token counts, without caching or tools. For the same text, the new tokenizer typically counts roughly 30% more tokens than Haiku 4.5. Recount your prompts. The Batch API offers a further 50% discount on input and output.
Model changes and how to access it
The Haiku 5.5 documentation specifies a one-million-token context window, up to 128,000 output tokens, and text and image input with text output. Its knowledge cutoff is June 2026. Haiku 4.5 had 200,000 context tokens and 64,000 output tokens. Message Batches has a beta option for up to 300,000 output tokens.
The Claude Platform ID is claude-haiku-5-5. Anthropic announces availability across its platforms and providers, including AWS, Google Cloud and Microsoft Azure. The documentation distinguishes provider IDs; Amazon Bedrock uses anthropic.claude-haiku-5-5.
For the first time in Haiku, you can adjust effort with low, medium, high, xhigh and max. Adaptive thinking is on by default. The claim that this is the fastest Claude refers to the models’ standard speed: the announcement notes that Opus in Fast Mode can run faster.
If you already have a Haiku 4.5 integration
The migration guide includes changes that can produce errors if you only swap the model name:
- Replace the manual thinking budget with adaptive thinking and adjust
effort. - Omit sampling parameters such as
temperature,top_pandtop_k; end the input conversation with a user turn. - Read response blocks by
type, because thinking may precede text. Handlestop_reason: "refusal". - For computer use on the Claude API and Google Cloud, move to the new toolset. Keep earlier turns unchanged and replay unmodified thinking blocks through the account that produced them.
The Python and TypeScript SDKs add beta support for browser use and computer use. Priority Tier is unavailable for Haiku 5.5 according to the migration guide.
Cheaper Sonnet caching; API credits for Max and Team
Sonnet 5.5 cache reads drop from $0.20 to $0.10 per million tokens. Anthropic estimates around 20% lower cost for agentic work where caching accounts for a large share of consumption. Input and output rates stay the same. Our Sonnet 5.5 launch story retains the cache price announced then; this is the October 7 update.
The monthly Claude Platform credits are rolling out over these days:
| Plan | Monthly credit |
|---|---|
| Max 5x | $100 |
| Max 20x | $200 |
| Team Standard | $20per seat |
| Team Premium | $100per seat |
Team pools the credits into one balance capped at $500. You must have spent seven days on an eligible plan and link a Console organization; credits expire at the cycle’s end and do not increase chat limits or cover interactive Claude Code sessions. Pro, Free and Enterprise are ineligible.
What early customer tests show
Anthropic publishes reports from six companies. These are their customers’ own tasks, with conditions different from the benchmarks above:
- Asana: over 30% lower task-completion latency in its AI Teammates evaluation, compared with the model it was using.
- HubSpot: 92.8% on its CRM task suite, averaged across three runs.
- AlphaSense: 0.84 versus 0.76 with Haiku 4.5 on 400 document queries.
- Box: 11 points above Haiku 4.5 and about half the latency in early tests.
- Rogo: describes extracting a figure from a financial filing with Haiku as a subagent while a larger model builds a presentation.
- Cognition: reports 66.2 on FrontierCode with Devin Fusion, Opus as the lead and Haiku as the sidekick. That is a combined system, not Haiku’s standalone score.
See the full attributions in the announcement. The system card also describes safety improvements and cases of excessive refusal. Cybersecurity safeguards are more restrictive than Haiku 4.5’s, while permitting more defensive tasks than Sonnet 5.5’s.
To decide whether it fits your work, start with a scoped task you can already evaluate: classify requests, extract data from files or summarize a conversation history. Measure the result and token consumption at the effort level you plan to use. For long, complex coding work, the launch table retains a Sonnet advantage.
Source
Plan your next post in HeyMark.
Keep the idea, review the draft with your team, and see how it performed in the accounts you connected.