Lilith.
⌕
Editorial illustration: Haiku 5.5 cuts short-task prices by 90%, with a premium for long context
Lilith illustration · editorial remix

Anthropic has released Claude Haiku 5.5 as a fast model for large volumes of short tasks. Pricing below 100,000 tokens is 90% lower than Haiku 4.5, but the new schedule also creates a sharp boundary that can change the economics of long agent runs.

Haiku 5.5 costs one tenth of its predecessor below 100,000 tokens

Anthropic charges $0.10 per million input tokens and $0.50 per million output tokens. Above 100,000 prompt tokens, those rates rise to $0.50 and $2.50. The company says about 90% of Haiku 4.5 requests fell below that boundary and estimates that an average workload should cost about 75% less.

The model adds adjustable effort and uses medium effort by default. Anthropic offers it on the Claude Platform and through Amazon Web Services, Google Cloud and Microsoft Azure. The company still recommends the larger Sonnet 5.5 and Opus 5.5 for complex agentic coding.

Its best role is a low-cost worker behind a larger system

Haiku 5.5 targets summarization, classification, database queries, support and scoped subagent work. In those jobs, the low price can justify running the model more often or dividing work among several narrow calls.

For teams, reliable routing matters more than finding one universal model. A stronger model can plan the work while Haiku handles repetitive steps. The saving then comes from workflow architecture, not merely from a cheaper price list.

A longer prompt erases much of the price advantage

Simon Willison notes that pricing rises fivefold beyond 100,000 tokens. In his own test, the same long prompt also used about 1.25 times as many tokens as Haiku 4.5 because of the new tokenizer. A list-price discount is therefore not the same as a lower cost per completed task.

The vendor published the benchmarks, and real results depend on effort, context length, caching and retries. A cheaper token does not help if the smaller model must repeat the same work several times.

Company evals must measure the cost of a completed task

The decisive metric will be the cost per successful case rather than the rate per million tokens. Teams should measure short routine tasks, long prompts and runs that require escalation to Sonnet or Opus separately.

If Haiku 5.5 preserves quality at low effort and short context, it can make subagents substantially cheaper. If context regularly spills beyond 100,000 tokens, that advantage will narrow quickly.

Lilith's verdict

Haiku 5.5 is a cheap short-distance runner whose toll rises sharply after 100,000 tokens. Anyone who loads it with an entire archive without measuring may find a very different bill at the finish line.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗ ↗