For AI agents: the public content index is available at https://we0.ai/llms.txt, and the English article bundle is available at https://we0.ai/llms-full.txt.
For AI agents: the complete content index is available at https://we0.ai/llms.txt, the full English article bundle is available at https://we0.ai/llms-full.txt, and this page is available as Markdown at https://we0.ai/articles/claude-haiku-5-vs-gpt-6-luna-benchmarks-p-ec36c306.md.
Anthropic has released **Claude Haiku 5.5**, and this is not a routine small-model refresh.

Anthropic has released Claude Haiku 5.5, and this is not a routine small-model refresh.
The new Haiku moves far beyond Haiku 4.5 in computer use, coding, professional work, and tool-heavy agent tasks while cutting its headline API price dramatically.
On Anthropic’s own launch comparisons, Haiku 5.5 also beats GPT-6 Luna on several benchmark entries, including GDPval-AA v2.1, OSWorld 2.1, Terminal-Bench 4.0, FrontierCode 1.1, and Chartography.
That makes Haiku 5.5 a serious new option for high-volume workloads such as:
But the launch also comes with important caveats.
The pricing changes at the 100K-prompt boundary, the tokenizer uses more tokens for the same text, and upgrading an existing Haiku 4.5 integration requires several API changes.

Anthropic positions Haiku 5.5 as its cheapest, fastest, and most capable small model to date.
The official benchmark table makes the intended competitor obvious: GPT-6 Luna is included directly beside Haiku 5.5.

The launch results include:
| Benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| GDPval-AA v2.1 | 1620 | 735 | 1437 | 1840 |
| AA-Briefcase v1.1 | 1578 | 614 | 1336 | 1824 |
| OSWorld 2.1, offline subset | 72.4% | 15.7% | 48.9% | 83.9% |
| Humanity’s Last Exam, no tools | 45.9% | 10.2% | — | 56.9% |
| Humanity’s Last Exam, with tools | 57.4% | 18.7% | — | 64.5% |
| Terminal-Bench 4.0 | 39.2% | 0.0% | 16.4% | 70.6% |
| FrontierCode 1.1 (Main) | 46.4% | — | 42.4% | 52.1% |
| Chartography, no tools | 46.4% | 6.4% | 29.1% | 61.6% |
The takeaway is not that Haiku 5.5 suddenly replaces every larger model.
It is that the smallest Claude tier has become much more capable at the kinds of tasks that used to require a larger model.
For prompts up to 100K tokens, Anthropic charges:
Haiku 5.5 input: $0.10 / 1M tokens
Haiku 5.5 output: $0.50 / 1M tokens
Cache read: $0.01 / 1M tokens
Cache write: $0.125 / 1M tokens
OpenAI’s standard GPT-6 Luna pricing is currently:
GPT-6 Luna input: $0.10 / 1M tokens
GPT-6 Luna cached input: $0.01 / 1M tokens
GPT-6 Luna cache write: $0.125 / 1M tokens
GPT-6 Luna output: $0.50 / 1M tokens
So for short-context standard API work, the two models line up almost exactly on the headline token rates.

Haiku 5.5 uses two pricing tiers.
| Token type | Prompt ≤100K | Prompt >100K |
|---|---|---|
| Cache read | $0.01/M | $0.05/M |
| Cache write | $0.125/M | $0.625/M |
| Input | $0.10/M | $0.50/M |
| Output | $0.50/M | $2.50/M |
Anthropic says prompts up to 100K tokens accounted for around 90% of Haiku 4.5 requests, which is why the lower tier covers the majority of expected traffic.
For prompts above 100K, Haiku 5.5 is still cheaper than Haiku 4.5, but the discount is smaller.
The source article compares Haiku 5.5 with DeepSeek-V4.1-Flash.
DeepSeek’s official pricing currently uses peak and off-peak rates.
For cache-miss input and output:
| Model | Input / 1M | Output / 1M |
|---|---|---|
| Haiku 5.5, ≤100K prompt | $0.10 | $0.50 |
| DeepSeek-V4.1-Flash, off-peak | $0.15 | $0.60 |
| DeepSeek-V4.1-Flash, peak | $0.30 | $1.20 |
DeepSeek remains cheaper on cache-hit input, especially off-peak, so the total cost still depends on workload structure.

The previous Haiku generation was not competitive on demanding computer-use tasks.
Haiku 5.5 changes that dramatically.
Anthropic’s OSWorld 2.1 offline-subset result rises from 15.7% for Haiku 4.5 to 72.4% for Haiku 5.5 at max effort.
The source article also performed a practical test: it asked Haiku 5.5 to add GPT-6 and DeepSeek pricing data into the Haiku launch-page table.
Instead of manually editing individual HTML elements through the UI, the model opened the browser console and injected the changes with script.

That behavior is useful because computer-use agents do not always need to imitate human mouse movement literally.
If a browser environment exposes a more direct execution path, using script can be faster and more reliable than clicking every cell.
Haiku 5.5 is the first Haiku-class model with an adjustable effort control.
The model supports a spectrum from lower-cost, faster behavior to more expensive, more deliberate reasoning.
Anthropic’s launch chart shows a clear accuracy-versus-cost curve on OSWorld.

The source highlights two points from that curve:
The exact cost of a production task will vary with prompt length, output length, caching, and tool use, but the overall pattern is important.
You can now tune Haiku based on the job instead of treating “small model” as one fixed capability level.
GDPval-AA v2.1 measures real-world professional work across 44 occupations.

At max effort, Anthropic reports an Elo of 1620 for Haiku 5.5, compared with 735 for Haiku 4.5.
Again, Sonnet 5.5 remains stronger, but the distance between “small” and “professional-useful” has narrowed sharply.
Haiku 5.5 introduces a newer tokenizer shared with newer Claude generations.
This matters because the same text does not produce the same token count as it did on Haiku 4.5.
Anthropic’s migration guide says the same input text produces approximately 30% more tokens on Haiku 5.5 than on Haiku 4.5, although the exact increase depends on the content.
The impact can be larger or smaller depending on the workload, with code, tables, and non-English text especially worth re-measuring.
That means this comparison is too simple:
Old input price: $1.00
New input price: $0.10
Therefore every task is exactly 90% cheaper
The better calculation is:
actual task cost
= new token count
× new token price
+ cache cost
+ output cost
+ tool cost
Anthropic itself says that after accounting for tokenizer changes, Haiku 5.5 costs around 75% less on average than Haiku 4.5 rather than a blanket 90% reduction for every task.
Artificial Analysis also shows why token efficiency matters when comparing inexpensive models.

A model can have a low token price while still using many more output tokens to finish a task.
For production systems, cost per completed task is more useful than price per million tokens alone.
The source article makes this point clearly: the “small cup” is still the small cup.
Haiku 5.5 is dramatically stronger than Haiku 4.5, but complex autonomous coding remains a larger-model job.
Terminal-Bench 4.0 shows the gap:
Haiku 5.5: 39.2%
Sonnet 5.5: 70.6%
Anthropic explicitly recommends Sonnet 5.5 or Opus 5.5 for complex agentic coding.
Haiku 5.5 is better suited to tasks that are already well-scoped:
Cognition is already using this structure in Devin Fusion.
The launch page says Opus 5.5 can act as the lead while Haiku 5.5 works as the sidekick.
That combination reached a 66.2 FrontierCode score in Cognition’s reported setup while reducing cost and latency.

The pattern is straightforward:
Opus / Sonnet
→ planning, decomposition, difficult judgment
Haiku 5.5
→ narrow subtasks, parallel execution, quick lookups, summaries
This is probably a more realistic way to use Haiku 5.5 than trying to replace a larger model everywhere.
If an application already uses Haiku 4.5, Anthropic warns that several changes can break existing requests.
The official model ID is now:
claude-haiku-5-5
But the migration has several additional requirements.
A Haiku 4.5 request like this is no longer valid:
{
"thinking": {
"type": "enabled",
"budget_tokens": 8000
}
}
Describe your idea once, and We0 AI can generate a showcase site, pages, and CMS, then help you attract customers and traffic after launch.
One complete project generation for free registration
Best for trying one complete generation flow and seeing a first project draft quickly.
Haiku 5.5 uses adaptive thinking and effort controls instead:
{
"thinking": {
"type": "adaptive"
},
"output_config": {
"effort": "medium"
}
}
A request that sends the old budget_tokens format returns a 400 error.
Adaptive thinking is also enabled by default, so applications should not assume the first response content block is always normal answer text.
Parse blocks by type.
max_tokensBecause the new tokenizer counts the same text as more tokens, old max_tokens values may no longer leave enough room for equivalent output.
Thinking tokens also count against the output budget.
Recount prompts using claude-haiku-5-5 rather than reusing Haiku 4.5 token estimates.
Haiku 5.5 no longer supports arbitrary sampling settings in the way Haiku 4.5 did.
Anthropic recommends omitting:
temperature
top_p
top_k
Non-default values can produce 400 errors, and top_k is not accepted.
Applications that used sampling parameters for creative variation should move that control into prompting or another application-level strategy.
Haiku 4.5 could continue from a final assistant message in some configurations.
Haiku 5.5 rejects assistant prefill.
If you previously used a prefix such as:
{"result":
to force JSON output, migrate to structured outputs or a tool schema instead.
For Claude API and Google Cloud integrations, the old computer-use declaration:
computer_20250124
must move to:
computer_toolset_20260801
The new toolset has a different event structure, so the agent loop must dispatch the returned tool-use blocks correctly and preserve toolset_name in results.
Haiku 5.5 also supports Anthropic’s browser-use toolset for webpage tasks on supported platforms.
Before deploying Haiku 5.5, check at least these items:
claude-haiku-5-5.budget_tokens thinking with adaptive thinking and effort.type rather than assuming block zero is answer text.computer_toolset_20260801 where required.refusal stop reason.Anthropic also provides an automated Claude Code migration command:
/claude-api migrate this project to claude-haiku-5-5
The tool can apply model-ID changes and detect several incompatible request patterns, but manual verification is still recommended.
Haiku 5.5 was not the only pricing change Anthropic announced.
Sonnet 5.5 cache-read pricing dropped from:
$0.20 / 1M tokens
to:
$0.10 / 1M tokens
Anthropic says cache reads make up a large share of token consumption in many agent workflows, so this change reduces the cost of most Sonnet 5.5 agentic tasks by roughly 20%.

This makes a mixed-model architecture more attractive:
Sonnet 5.5
→ planning and harder execution
Haiku 5.5
→ inexpensive high-volume subtasks
Anthropic also introduced monthly Claude Platform credits for eligible Max and Team subscribers.
The current official credit amounts are:
| Plan | Monthly API credit |
|---|---|
| Max 5x | $100 |
| Max 20x | $200 |
| Team Standard seat | $20 per seat |
| Team Premium seat | $100 per seat |
| Team pooled maximum | $500 |

The credits can be used for Claude Platform workloads such as:
They do not increase the user’s Claude or Claude Code subscription usage limits, and they do not cover normal interactive Claude Code sessions when those sessions are billed through the subscription plan.
Credits are linked to one Claude Console organization and do not roll over indefinitely.
The source closes with a lighter comparison.
While Anthropic announced cheaper models, cheaper Sonnet cache reads, and monthly API credits, OpenAI’s Codex and ChatGPT Work lead Thibault Sottiaux announced another banked reset for paid accounts while celebrating 40 million active users across Codex and ChatGPT Work.

A banked reset is different from a permanent price cut.
It is an additional usage reset that eligible users can save and redeem rather than a change to the model’s ongoing token pricing.
The comparison is therefore more rhetorical than technical:
Anthropic launch:
new small model + lower cache price + monthly API credits
OpenAI update:
additional banked usage reset for paid Codex / Work accounts
Both are attempts to give heavy users more room to experiment, but they operate through different mechanisms.
The launch makes Haiku much more useful, but the ideal role remains relatively clear.
Use Haiku 5.5 for:
Prefer a larger Claude model for:
The cheapest model is not always the cheapest way to finish a job.
For production systems, measure:
cost per successful task
=
input + cache + output + tool use + retries + failures
That metric will tell you more than a model’s headline input price.
Claude Haiku 5.5 is Anthropic’s latest small Claude model, released for high-volume, cost-sensitive, and latency-sensitive workloads. It is also the first Haiku model with adjustable effort and is substantially stronger than Haiku 4.5 on coding, computer use, and professional-work evaluations.
For prompts up to 100K tokens, it costs $0.10 per million input tokens and $0.50 per million output tokens. For prompts over 100K, the rates rise to $0.50 input and $2.50 output per million tokens, with separate cache-read and cache-write prices.
For standard short-context token pricing, their headline input, cached-input, cache-write, and output rates are effectively aligned. Haiku 5.5’s price increases above the 100K prompt threshold, so long-context workloads should be compared separately.
No. Anthropic’s own Terminal-Bench 4.0 results show Sonnet 5.5 at 70.6% versus Haiku 5.5 at 39.2%. Haiku is better positioned for narrow, repeatable, high-volume tasks and subagent work, while Sonnet and Opus remain better for complex autonomous coding.
Haiku 5.5 uses a newer tokenizer. Anthropic’s migration guide says the same text produces about 30% more tokens on average than Haiku 4.5, although the exact increase depends on content.
Yes. Important changes include adaptive thinking instead of manual budget_tokens, revised sampling-parameter behavior, removal of assistant prefill, a new computer-use toolset, and the need to parse response blocks by type. Anthropic provides an official migration guide and a Claude Code migration command.
On the Claude API, the model ID is claude-haiku-5-5. Anthropic documents corresponding IDs for Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS.
Eligible Max and Team subscribers can claim monthly Claude Platform credits. Max 5x receives $100, Max 20x receives $200, and Team credits are calculated by seat type and pooled up to $500 per month.
Claude Haiku 5.5 is a major upgrade to Anthropic’s small-model tier. It beats Haiku 4.5 by a wide margin on computer use, professional work, coding, and reasoning, while matching GPT-6 Luna’s short-context headline API pricing and beating Luna on several of Anthropic’s published benchmark comparisons.
The lower price is real, but production teams should not ignore the new tokenizer or the 100K-prompt pricing boundary. The same text can consume about 30% more tokens than on Haiku 4.5, and long prompts use the higher pricing tier.
The migration also includes genuine breaking changes: adaptive thinking replaces manual thinking budgets, custom sampling controls are restricted, assistant prefill is removed, and computer-use integrations need the newer toolset. Existing Haiku 4.5 applications should follow the migration guide rather than changing only the model name.
Haiku 5.5 is best understood as a fast, inexpensive execution model for scale and subagents—not as a universal replacement for Sonnet or Opus on complex long-horizon work.
Start from one sentence and have a complete website in minutes.