Anthropic has launched Claude Haiku 5.5, its fastest and least expensive small AI model, bringing substantially lower API pricing and stronger computer-use capabilities to developers building autonomous AI agents.
Released on October 7, 2026, the model costs as little as $0.10 per million input tokens, a 90% reduction in the per-token price compared with Claude Haiku 4.5 for prompts containing no more than 100,000 tokens.
The performance improvements are equally significant. According to Anthropic’s official announcement, Haiku 5.5 achieved 72.4% on the offline subset of OSWorld 2.1, a benchmark measuring an AI system’s ability to operate computers and complete multi-step tasks.
That score is close to the 72.36% human baseline published for the original OSWorld benchmark, although the evaluations are not identical.
For developers, the combination of stronger capabilities and lower costs could make smaller AI models particularly attractive for automating repetitive work.
Key Takeaways:
- Claude Haiku 5.5 starts at $0.10 per million input tokens, compared with $1 for Haiku 4.5.
- The model scored 72.4% on OSWorld 2.1’s offline subset, according to Anthropic’s evaluation.
- Real-world costs are approximately 75% lower on average, Anthropic estimates, rather than a universal 90% reduction.
- A 1-million-token context window and adjustable reasoning effort support longer workflows and specialized AI agents.
- Haiku 5.5 is available through Anthropic’s API and major cloud platforms, including AWS, Google Cloud, and Microsoft Azure.
Claude Haiku 5.5 Pricing: The Full Breakdown
Anthropic has introduced a two-tier pricing structure that depends on prompt length.
Requests containing up to 100,000 input tokens qualify for the lowest rates. Longer prompts use a higher pricing tier, although both are less expensive per token than Haiku 4.5.
The following table compares the official API prices across Haiku 5.5, its predecessor, and Sonnet 5.5.
All prices are in US dollars per 1 million tokens.
| Pricing Category | Haiku 5.5 (≤100K) | Haiku 5.5 (>100K) | Haiku 4.5 | Sonnet 5.5 |
|---|---|---|---|---|
| Input tokens | $0.10 | $0.50 | $1.00 | $2.00 |
| Output tokens | $0.50 | $2.50 | $5.00 | $10.00 |
| Cache writes (5 minutes) | $0.125 | $0.625 | $1.25 | $2.50 |
| Cache writes (1 hour) | $0.20 | $1.00 | $2.00 | $4.00 |
| Cache reads | $0.01 | $0.05 | $0.10 | $0.10 |
Source: Anthropic’s official model pricing documentation.
The distinction between input and output tokens matters. Input tokens represent information sent to the model, while output tokens represent the response it generates.
Prompt caching can further reduce expenses by reusing previously processed information instead of repeatedly processing the same content.
Anthropic also offers a 50% Batch API discount on input and output tokens, making workloads that do not require immediate responses more economical.
There is one important qualification to the headline price reduction.
Haiku 5.5 uses a newer tokenizer that can produce approximately 30% more tokens from the same text than Haiku 4.5. Consequently, a 90% reduction in the published per-token rate does not guarantee identical savings for every application.
Anthropic estimates the average effective cost reduction at approximately 75%.
Computer-Use Performance Reaches a Major Milestone
Computer use is one of Haiku 5.5’s most important improvements.
Instead of merely describing how to complete a task, a computer-use agent can potentially interact with software interfaces, navigate applications, and perform operations on a virtual desktop.
Anthropic reports that Haiku 5.5 scored 72.4% on the OSWorld 2.1 offline subset, compared with 15.7% for Haiku 4.5.
The larger Sonnet 5.5 achieved 83.9% on the same reported evaluation.
| AI Model | OSWorld 2.1 Offline Subset |
|---|---|
| Claude Sonnet 5.5 | 83.9% |
| Claude Haiku 5.5 | 72.4% |
| GPT-6 Luna | 48.9% |
| Claude Haiku 4.5 | 15.7% |
Source: Anthropic’s published benchmark results. The figures are company-reported and should not be treated as an independent audit.
The original OSWorld research project established a human baseline above 72% for computer interaction tasks.
However, matching that historical figure numerically does not establish that Haiku 5.5 performs at human level across all desktop activities. Benchmark versions, task subsets, and testing conditions matter.
The improvement nevertheless suggests that smaller models are becoming more capable of handling practical software automation.
Why Haiku 5.5 Could Change the Economics of AI Agents
The financial advantage becomes especially important when developers combine multiple AI models into a coordinated system.
Imagine an AI coding assistant managing a complicated software project.
A larger model, such as Claude Opus 5.5, could analyze the problem and decide which changes are needed. A smaller model could then inspect files, retrieve relevant information, and complete narrower tasks.
These smaller models are often called sub-agents.
Using an expensive model for every routine operation can significantly increase the cost of an application. Delegating straightforward work to cheaper agents offers a way to reduce that expense without removing more capable models from decisions that require deeper reasoning.
Anthropic highlighted an example involving Cognition’s Devin Fusion.
According to results shared through Anthropic’s announcement, Devin Fusion achieved a FrontierCode 1.1 score of 66.2 using Haiku 5.5 alongside Opus 5.5, while reducing cost and latency.
These are customer-reported results, not proof that every multi-model application will achieve similar improvements.
Other early customers, including Asana, HubSpot, and Box, also reported favorable results involving speed, analytical tasks, or application-specific evaluations.
The larger implication is that increasingly capable small models may allow developers to reserve premium AI processing for tasks where it delivers the greatest value.
A Larger Context Window and Adjustable Reasoning Effort
Haiku 5.5 offers more than lower pricing.
According to its official technical documentation, the model supports a context window of up to 1 million tokens and standard maximum outputs of 128,000 tokens.
That larger working context can help applications process extensive documents, software repositories, and long-running conversations.
Haiku 5.5 also introduces adaptive thinking with an adjustable effort parameter.
Developers can reduce reasoning effort for straightforward requests or increase it for more demanding tasks.
This flexibility matters because not every operation requires the same amount of computation.
For instance, classifying an email and solving a complicated programming problem are fundamentally different workloads.
By adjusting effort and selecting the appropriate model, developers can better balance speed, accuracy, and cost.
Anthropic has also introduced beta support for computer-use and browser-use capabilities through its Python and TypeScript software development kits.
These tools are intended to make it easier to build applications that interact with software interfaces rather than relying exclusively on traditional APIs.
Availability, Security, and What Comes Next
Claude Haiku 5.5 is available through the Claude API using the model identifier claude-haiku-5-5.
Anthropic also lists availability through Amazon Bedrock, Google Cloud, and Microsoft Foundry.
The launch includes a separate pricing adjustment for Claude Sonnet 5.5: cached input reads have dropped from $0.20 to $0.10 per million tokens.
Anthropic estimates that this reduces the cost of typical agentic workloads using Sonnet 5.5 by approximately 20%.
Security remains an important consideration.
Anthropic says Haiku 5.5 performed better than its predecessor in internal alignment evaluations and shows less willingness to assist with misuse. These are company assessments, and real-world safety also depends on the applications, permissions, and tools connected to the model.
As AI agents gain access to browsers, documents, and external services, developers will need safeguards against unauthorized actions, unreliable results, and malicious instructions embedded in external content.
The bigger picture: Haiku 5.5 is not a universal replacement for larger AI models. Its importance lies in making everyday AI operations substantially less expensive while improving the capabilities available to smaller agents.
For businesses running large volumes of AI requests, that combination could prove more consequential than another incremental improvement in chatbot performance.
Explore more developments in artificial intelligence, AI models, and automation at Tech News Home.
Frequently Asked Questions About Claude Haiku 5.5
How much does Claude Haiku 5.5 cost?
Claude Haiku 5.5 starts at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. Longer prompts cost $0.50 and $2.50 per million tokens, respectively. Prompt caching and batch processing can further change overall API costs.
Is Claude Haiku 5.5 better than Claude Haiku 4.5?
Anthropic reports substantial improvements across computer-use, reasoning, and other benchmarks. Haiku 5.5 also adds adjustable reasoning effort, a larger context window, and lower API rates. Actual improvements depend on the application and workload.
Can Claude Haiku 5.5 control a computer?
Yes. Claude Haiku 5.5 can support computer-use workflows through compatible tools and integrations. It can help AI agents navigate software interfaces and perform tasks, although reliable automation requires appropriate permissions, tools, and testing.
Official Sources and Further Reading
- Anthropic — Introducing Claude Haiku 5.5 — Official October 7, 2026 announcement covering performance, pricing changes, early customer evaluations, and availability.
- Claude Platform — Claude Haiku 5.5 Documentation — Technical specifications, complete input and output pricing, cache rates, Batch API discount, context limits, and model availability.
- Claude Platform — Claude Sonnet 5.5 Documentation — Official comparison pricing and technical specifications for Anthropic’s larger model.
- OSWorld — Computer-Use Benchmark Research — Primary research explaining desktop-agent evaluations, task environments, and the original human baseline.