Anthropic Launches Haiku 5.5: Its Cheapest and Fastest Claude Model Yet
Anthropic has launched Claude Haiku 5.5, a lower-cost, faster small model priced at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. The company positions Haiku 5.5 for high-volume and latency-sensitive tasks such as document summarization and live customer support and reports performance gains on several agent and professional-work benchmarks versus earlier Haiku and competitor small models.

Why It Matters
The new pricing brings Haiku 5.5 roughly in line with rival small-model pricing and — according to Anthropic — cuts average running costs for users by about 75%, which could materially lower expenses for applications that run many repetitive, high-volume prompts. At the same time, the model introduces features aimed at production use, including an adjustable effort setting and reduced cache-read pricing.
Key Facts
- Release: Claude Haiku 5.5 released Wednesday (named claude-haiku-5-5 on cloud marketplaces)
- Pricing (<=100k-token prompts): $0.10 per million input tokens; $0.50 per million output tokens
- Previous Haiku pricing: Haiku 4.5 charged $1 per million input and $5 per million output tokens
- Anthropic's estimated average savings: About 75% average saving vs Haiku 4.5
- Benchmarks — OSWorld 2.1: Haiku 5.5: 72.4% vs OpenAI GPT-6 Luna: 48.9%
Anthropic introduced Claude Haiku 5.5 as its most economical and latency-optimized small model to date, targeting high-volume tasks like document summarization, database queries and live customer-support interactions. The company set input token pricing at $0.10 per million and output token pricing at $0.50 per million for prompts up to 100,000 tokens; prompts above that threshold receive a 50% discount. Anthropic notes that roughly 90% of requests to the prior Haiku model were below 100,000 tokens and, due to tokenization differences, estimates average cost reductions near 75% compared with Haiku 4.5.
On standard agent and professional-work benchmarks, Anthropic reports Haiku 5.5 outperformed OpenAI’s GPT-6 Luna and its own prior Haiku 4.5 in several measures. Haiku 5.5 scored 72.4% on OSWorld 2.1 versus Luna’s 48.9%. On Terminal-Bench 4.0 — which assesses an agent’s ability to complete tasks by issuing commands — Haiku 5.5 achieved 39.2%, compared with 16.4% for Luna and 0% for Haiku 4.5; Anthropic’s larger Sonnet 5.5 scored 70.6% on that coding-style test. On GDPval-AA v2.1, a multi-occupation professional-work Elo-style metric, Haiku 5.5 posted 1620 against Luna’s 1437 and Haiku 4.5’s 735.
Anthropic also added operational features aimed at production deployments. Haiku 5.5 is the first Haiku model to include an adjustable effort setting that lets developers trade lower cost for higher compute effort and presumably smarter responses. The company cut Sonnet 5.5’s cache-read price to $0.10 per million tokens and is rolling out monthly API credits for subscribers: $100 for Max 5x, $200 for Max 20x, and up to $500 pooled for Team plans.
Haiku 5.5 arrives shortly after Anthropic’s other Claude 5.5 releases: it followed Opus 5.5 by 15 days (Opus launched Sept. 22) and came nine days after Sonnet 5.5 (Sept. 28). The model is available now through Anthropic’s website and on Amazon Web Services, Google Cloud and Microsoft Azure under the name claude-haiku-5-5. Anthropic cautions users that while the model is fast, it can still produce incorrect answers and should not be trusted blindly.
Keep Reading
Microsoft and Nvidia are teaming up on a supercharged AI laptop

Google Lets You Build a Video Game by Typing a Description

Hackers Used AI Agents to Raid a Megachurch's Database, Exposing 850,000 Members
