HeyMark
Content

News Claude

Anthropic launches Claude Haiku 5.5: benchmarks and new pricing

7 min read
Anthropic's official launch image with the Claude Haiku 5.5 name and three blocks of color.

Claude Haiku 5.5 is available now. Anthropic released it on October 7 for high-volume work: extracting data, classifying requests, summarizing documents and handling scoped tasks within agents. API pricing starts at $0.10 per million input tokens and $0.50 per million output tokens, for prompts up to 100,000 tokens.

The official announcement also brings changes for people using other Claude models: Sonnet 5.5 cache reads at half the price, and new monthly API credits for Max and Team. This story’s cover image is from Anthropic.

Haiku 5.5 vs. Haiku 4.5, GPT-6 Luna and Sonnet 5.5

Choose a test to compare the launch results. GDPval-AA and AA-Briefcase use Elo ratings; the other tests show percentages. Each measures different tasks: compare models within the same benchmark.

07.10.2026 · Anthropic

Compare by task

GDPval-AA v2.1

Higher is better

Professional deliverables across 44 occupations. Elo compares results; it is not an accuracy percentage.

Haiku 5.5
1620
Haiku 4.5
735
GPT-6 Luna
1437
Sonnet 5.5Reference
1840

Independently evaluated by Artificial Analysis. Haiku 5.5 scores 1620 at max effort; at medium, the default, it scores 1277.

AA-Briefcase v1.1

Higher is better

Long-running professional projects with linked tasks and many source files. Elo score.

Haiku 5.5
1578
Haiku 4.5
614
GPT-6 Luna
1336
Sonnet 5.5Reference
1824

Independently evaluated by Artificial Analysis. Haiku 5.5 scores 1372 at medium effort and 1578 at max.

OSWorld 2.1

Higher is better

Computer use, offline subset. These are not OSWorld 2.0 results or results for the full task set.

Haiku 5.5
72.4%
Haiku 4.5
15.7%
GPT-6 Luna
48.9%
Sonnet 5.5Reference
83.9%

Anthropic evaluated GPT-6 Luna on the same 82 tasks. The system card distinguishes partial-credit scoring from passing every checkpoint in a task.

Humanity’s Last Exam · no tools

Higher is better

Multidisciplinary academic knowledge and reasoning, without tools.

Haiku 5.5
45.9%
Haiku 4.5
10.2%
GPT-6 Luna
—
No reported result
Sonnet 5.5Reference
56.9%

The table does not publish a GPT-6 Luna result for this test. A dash means no reported result.

Humanity’s Last Exam · with tools

Higher is better

The same evaluation family with tools available. This is a different configuration from the previous chart.

Haiku 5.5
57.4%
Haiku 4.5
18.7%
GPT-6 Luna
—
No reported result
Sonnet 5.5Reference
64.5%

Do not combine results with and without tools into one score. There is no published GPT-6 Luna result here.

Terminal-Bench 4.0

Higher is better

Agentic coding tasks in a terminal.

Haiku 5.5
39.2%
Haiku 4.5
0.0%
GPT-6 Luna
16.4%
Sonnet 5.5Reference
70.6%

Haiku 4.5's 0.0% is the reported result. Haiku 5.5 improves, but Sonnet 5.5 retains a substantial lead on this test.

FrontierCode 1.1 · Main

Higher is better

The 100 hardest tasks in Cognition's coding evaluation.

Haiku 5.5
46.4%
Haiku 4.5
—
No reported result
GPT-6 Luna
42.4%
Sonnet 5.5xhigh
52.1%

The launch table uses Sonnet 5.5 at xhigh: 52.1%. The system card also reports 46.2% at max. Haiku 5.5 uses max; no Haiku 4.5 result is listed.

Chartography

Higher is better

Visual reasoning about charts, without tools.

Haiku 5.5
46.4%
Haiku 4.5
6.4%
GPT-6 Luna
29.1%
Sonnet 5.5Reference
61.6%

This evaluation does not measure copy quality or campaign performance.

View the full table
Launch results, with units
BenchmarkHaiku 5.5Haiku 4.5GPT-6 LunaSonnet 5.5
GDPval-AA v2.1 (Elo)162073514371840
AA-Briefcase v1.1 (Elo)157861413361824
OSWorld 2.1 (%)72.4%15.7%48.9%83.9%
Humanity’s Last Exam · no tools (%)45.9%10.2%—56.9%
Humanity’s Last Exam · with tools (%)57.4%18.7%—64.5%
Terminal-Bench 4.0 (%)39.2%0.0%16.4%70.6%
FrontierCode 1.1 · Main (%)46.4%—42.4%52.1% (xhigh)
Chartography (%)46.4%6.4%29.1%61.6%

The improvement over Haiku 4.5 spans very different tasks. On OSWorld 2.1 offline, its score rises from 15.7% to 72.4%. On Terminal-Bench 4.0 it reaches 39.2%, against the 16.4% reported for GPT-6 Luna; Sonnet 5.5 scores 70.6%. That gap matters when choosing an agent for complex coding work.

The configuration changes the comparison

The system card, section 8, explains the environments and reasoning effort. Haiku 5.5 is generally evaluated at max, while the API defaults to medium. On GDPval-AA it scores 1620 Elo at max and 1277 at medium; at medium it uses about a tenth of the output tokens. A low token price and a benchmark score alone therefore cannot tell you what your task will cost.

There is another visible difference in FrontierCode: the announcement lists Sonnet at 52.1% with xhigh. The system card also reports 46.2% at max. Our chart retains the launch value and its setting. Dashes mean no reported result; Haiku 4.5’s 0.0% on Terminal-Bench is an actual published value.

Haiku 5.5 API pricing: two prompt-length tiers

The pricing documentation sets the tier by prompt length. Above 100,000 tokens, the request’s input, output and caching rates change.

API rates per million tokens · USD
Standard-rate comparison, including both Haiku 5.5 prompt tiers
UsageHaiku 5.5Haiku 4.5Sonnet 5.5
Up to 100,000prompt tokensOver 100,000prompt tokens
Input$0.10$0.50$1.00$2.00
Output$0.50$2.50$5.00$10.00
Cache read$0.01$0.05$0.10$0.10
Cache write · 5 min$0.125$0.625$1.25$2.50
Cache write · 1 h$0.20$1.00$2.00$4.00

Scroll across to compare all four rates.

Per-token pricing is 90% lower than Haiku 4.5 in the short-prompt tier and 50% lower in the long tier. Anthropic estimates about 75% lower cost per task on average, accounting for its request mix and changes in token consumption. That average does not promise identical savings for every integration.

One message, two types of tokens

You write a message and Claude replies. The API counts each text separately: input for your message, output for the reply. A token is a piece of text; a word can take more than one.

Write a short reply thanking a customer for their comment.

Thanks for sharing your experience! We're glad it was helpful.

How can I help you today?
Local simulation using the Claude interface.
Your message · input80 tokens
The reply · output40 tokens

Cost per request

Prompt up to 100,000 tokens

Haiku 5.5
$0.000028$0.028 per 1,000 identical requests
Haiku 4.5
$0.00028$0.280 per 1,000 identical requests
Sonnet 5.5
$0.00056$0.560 per 1,000 identical requests

The example uses 80 input tokens and 40 output tokens. Text and counts are illustrative; each model may count and respond differently. Prices use API rates; they do not calculate your Claude subscription bill.

Standard API rates in USD, excluding caching and tool charges. Thinking tokens also count as output. Compares token volumes; the cost of a task depends on what each model consumes.

Official rates ↗

The calculator lets you try the jump from 100,000 to 100,001 tokens. It compares equal token counts, without caching or tools. For the same text, the new tokenizer typically counts roughly 30% more tokens than Haiku 4.5. Recount your prompts. The Batch API offers a further 50% discount on input and output.

Model changes and how to access it

The Haiku 5.5 documentation specifies a one-million-token context window, up to 128,000 output tokens, and text and image input with text output. Its knowledge cutoff is June 2026. Haiku 4.5 had 200,000 context tokens and 64,000 output tokens. Message Batches has a beta option for up to 300,000 output tokens.

The Claude Platform ID is claude-haiku-5-5. Anthropic announces availability across its platforms and providers, including AWS, Google Cloud and Microsoft Azure. The documentation distinguishes provider IDs; Amazon Bedrock uses anthropic.claude-haiku-5-5.

For the first time in Haiku, you can adjust effort with low, medium, high, xhigh and max. Adaptive thinking is on by default. The claim that this is the fastest Claude refers to the models’ standard speed: the announcement notes that Opus in Fast Mode can run faster.

If you already have a Haiku 4.5 integration

The migration guide includes changes that can produce errors if you only swap the model name:

  • Replace the manual thinking budget with adaptive thinking and adjust effort.
  • Omit sampling parameters such as temperature, top_p and top_k; end the input conversation with a user turn.
  • Read response blocks by type, because thinking may precede text. Handle stop_reason: "refusal".
  • For computer use on the Claude API and Google Cloud, move to the new toolset. Keep earlier turns unchanged and replay unmodified thinking blocks through the account that produced them.

The Python and TypeScript SDKs add beta support for browser use and computer use. Priority Tier is unavailable for Haiku 5.5 according to the migration guide.

Cheaper Sonnet caching; API credits for Max and Team

Sonnet 5.5 cache reads drop from $0.20 to $0.10 per million tokens. Anthropic estimates around 20% lower cost for agentic work where caching accounts for a large share of consumption. Input and output rates stay the same. Our Sonnet 5.5 launch story retains the cache price announced then; this is the October 7 update.

The monthly Claude Platform credits are rolling out over these days:

Monthly API credits · USD
Monthly credit by Claude plan
PlanMonthly credit
Max 5x$100
Max 20x$200
Team Standard$20per seat
Team Premium$100per seat

Team pools the credits into one balance capped at $500. You must have spent seven days on an eligible plan and link a Console organization; credits expire at the cycle’s end and do not increase chat limits or cover interactive Claude Code sessions. Pro, Free and Enterprise are ineligible.

What early customer tests show

Anthropic publishes reports from six companies. These are their customers’ own tasks, with conditions different from the benchmarks above:

  • Asana: over 30% lower task-completion latency in its AI Teammates evaluation, compared with the model it was using.
  • HubSpot: 92.8% on its CRM task suite, averaged across three runs.
  • AlphaSense: 0.84 versus 0.76 with Haiku 4.5 on 400 document queries.
  • Box: 11 points above Haiku 4.5 and about half the latency in early tests.
  • Rogo: describes extracting a figure from a financial filing with Haiku as a subagent while a larger model builds a presentation.
  • Cognition: reports 66.2 on FrontierCode with Devin Fusion, Opus as the lead and Haiku as the sidekick. That is a combined system, not Haiku’s standalone score.

See the full attributions in the announcement. The system card also describes safety improvements and cases of excessive refusal. Cybersecurity safeguards are more restrictive than Haiku 4.5’s, while permitting more defensive tasks than Sonnet 5.5’s.

To decide whether it fits your work, start with a scoped task you can already evaluate: classify requests, extract data from files or summarize a conversation history. Measure the result and token consumption at the effort level you plan to use. For long, complex coding work, the launch table retains a Sonnet advantage.

Source

Anthropic · Introducing Claude Haiku 5.5

Plan your next post in HeyMark.

Keep the idea, review the draft with your team, and see how it performed in the accounts you connected.