FR
live
AI

Anthropic slashes Claude Haiku 5.5’s price tenfold and adds effort controls to its smallest model

On October 7, 2026, Anthropic launched Claude Haiku 5.5, its smallest model, at $0.10 per million input tokens — a 90% cut — with effort controls never seen before on this tier. If you run high-volume agentic tasks, this price reshapes the small-model decision.

A single small amber LED lit among a field of larger, dark LEDs on a dark circuit board.

October 7, 2026. Anthropic launched Claude Haiku 5.5, the first refresh of its smallest model in nearly a year. $0.10 per million input tokens, against $1 for the previous generation. 72.4% on the offline subset of OSWorld 2.1, against 15.7% for Haiku 4.5. Why it matters: the cheapest model in the lineup just became ten times cheaper and markedly more capable — and it is the third 5.5 model Anthropic has shipped in a month.

The price cut is not cosmetic

The price change is the headline, and it is structural. Haiku 4.5 cost $1 in and $5 out per million tokens, regardless of volume. Haiku 5.5 introduces size-based tiering: $0.10 / $0.50 for requests under 100,000 tokens, and $0.50 / $2.50 beyond that. Those are 90% and 50% cuts respectively, for an average saving around 75% by Anthropic’s accounting — once you factor in a revised tokenizer that uses slightly more tokens per task.

The split matters: Anthropic notes that about 90% of requests to Haiku 4.5 already fell in the cheaper tier. The cut therefore targets the heart of real traffic — the many short tasks that make up the bulk of agentic calls — rather than a marginal segment.

Effort controls reach the small model

Haiku 5.5 is the first Haiku model with effort controls, set to medium by default. A developer can now trade, on one model, between a minimal pass and a longer reasoning pass depending on the task. It is a direct transfer from the Sonnet and Opus families down the stack, and it changes how a small model gets used.

The logic echoes the decision models — the category where players like Jev are currently making noise — but applied to a cheap general-purpose model. Concretely, you can keep effort low for context compaction and database queries, and push it higher for tasks that demand reasoning, such as browser navigation or live customer support.

Unusual capability gains for a small model

Anthropic’s internal evaluations tell a rare story for a “small and efficient” model: a leap, not a refresh. On the offline subset of OSWorld 2.1, which measures computer use, Haiku 5.5 hits 72.4%, up from 15.7% for Haiku 4.5 — and ahead of GPT-6 Luna’s 48.9%, OpenAI’s small model.

The pattern repeats elsewhere. On GDPval-AA v2.1, a knowledge-work benchmark, Haiku 5.5 scores 1,620, against 735 for its predecessor and 1,437 for GPT-6 Luna. On Terminal-Bench 4.0, which tests complex multi-step command-line tasks, it goes from 0% to 39.2% — still far from Sonnet 5.5’s 70.6%, but on a model whose job is not to be the strongest.

Read these numbers with the usual caution: they are Anthropic’s own evaluations, and they compare Haiku 5.5 only to its own models and GPT-6 Luna. They still say something true — the line between “small model” and “capable model” moved in a year.

The competition is not asleep

The small-model space is exactly where price pressure is fiercest, and Haiku 5.5 is not alone there. Artificial Analysis credits Z.ai’s GLM-5.3-Flash with 1,647 on GDPval-AA v2.1 and 1,454 on AA-Briefcase v1.1, against 1,620 and 1,578 for Haiku 5.5 — neighboring scores at sometimes lower prices.

On price, the reference remains Alibaba’s Qwen3.7 Flash: $0.03 per million input tokens and $0.13 out up to 32,000 tokens, then $0.10 / $0.40 up to 256,000 tokens. On paper, that is still cheaper than Haiku 5.5. The decision therefore hinges less on absolute price than on the ecosystem: computer-use quality, guardrail robustness, and availability on the clouds the enterprise already uses.

Availability and side benefits

Haiku 5.5 is available on the Claude Platform, AWS, Google Cloud, and Microsoft Azure, reachable as claude-haiku-5-5. Anthropic is also adding beta computer-use and browser-use support to its Python and TypeScript SDKs.

In the same announcement, Anthropic halves Sonnet 5.5’s cache-read price, from $0.20 to $0.10 per million tokens — enough to make most agentic tasks about 20% cheaper. Monthly API credits are also coming to Max and Team plans (up to $200 for Max 20x, $500 pooled for Team). Finally, Haiku 5.5 carries tighter cybersecurity safeguards than Haiku 4.5, allowing a wider range of defensive work than Sonnet 5.5’s, while still blocking penetration testing.

The real cost of a high-volume agent

To judge the cut, a concrete number helps. Take an agent handling one million short requests a month, each around 2,000 input tokens and 500 output tokens — the order of magnitude of a context compaction or a database query. On Haiku 4.5, at $1 / $5 per million tokens, the monthly bill reaches $2,000 in and $2,500 out — $4,500. On Haiku 5.5, in the under-100,000-token tier, the same load drops to $200 in and $250 out — $450, ten times less.

The gap widens once effort controls enter the picture. A live customer-support agent can leave effort at medium and absorb the bulk of requests at low cost, then switch to high effort for the cases that demand reasoning. On Haiku 4.5, that dial did not exist: everything ran at the same price regardless of real need. The new granularity turns the small model into a budget lever — you pay for thinking only when you need it.

The remaining question is quality. Ten times cheaper is worth nothing if the model fails where you need it; that is why the jump from 15.7% to 72.4% on OSWorld 2.1 matters as much as the price. A model that costs ten times less and succeeds at four times more computer-use tasks is not a promotion — it is a category change.

Strategically, the timing is telling. Anthropic has now shipped Sonnet 5.5, Opus 5.5, and Haiku 5.5 in a single month, with Fable 5.5 still pending a longer safety review. The deliberate order — flagship, mid-tier, then the price-sensitive small model — reads as a single move: own the agentic workload at every price point before the competition locks in the high-volume bottom of the market. For OpenAI, the counterweight is GPT-6 Luna, its own small model — but at $0.10 per million tokens and a 72.4% OSWorld score, Haiku 5.5 now sets a high bar on both price and capability at once, a combination that is hard to undercut on either axis alone.

Verdict

If you push a high volume of short tasks — compaction, database queries, live support, browser use — Haiku 5.5 changes the equation: at $0.10 per million tokens, its computer-use leap makes it a serious alternative to GPT-6 Luna. If you optimize for pure price, benchmark GLM-5.3-Flash and Qwen3.7 Flash first — cheaper on paper — before deciding on the strength of Anthropic’s benchmarks: measure against your workload, not the press release. If you are already on Sonnet 5.5, take the cache-read cut first, which lowers your agentic costs by about 20% without changing models. The small model is no longer the lineup’s poor relation — it is now the economic workhorse, and the economics favor shipping more agent calls to it while reserving the larger models for the hard cases.

References

The cyber brief, every Tuesday

The flaws that matter and the patches to apply, in a ten-minute read.

No spam. One-click unsubscribe.
read next

On the same topic

Mistral opens Large 4, a 1.05-trillion-parameter model whose weights land at the end of October

On October 6, 2026, Mistral AI shipped a public preview of Mistral Large 4, nicknamed “Le Chonk”: a multimodal mixture-of-experts model of roughly 1.05 trillion parameters, 49 billion active per token, with open weights promised for October 27. Test it on your own workloads now, but withhold any judgment on the benchmarks until the weights and the license are actually published.

Reflection AI opens Beam, a 501B model that challenges Chinese open weights

On October 5, 2026, Reflection AI unveiled Beam, a 501-billion-parameter open-source model trained in eight weeks on NVIDIA hardware rented from SpaceX, and which approaches the performance of two-trillion-parameter models. For teams evaluating open-weight models, Beam shifts the cost-performance calculus, but its weights only land at the end of the month and still trail closed frontier models.

← Back to the feed

Type at least two characters.

↑ ↓ navigate ↵ open esc dismiss