Anthropic released Claude Haiku 5.5 overnight, and the pace has not eased. Last year’s small model, Haiku 4.5, cost US$1 in and US$5 out per million tokens. Haiku 5.5 is US$0.10 and US$0.50, a tenth of the price, and Anthropic calls it its cheapest, fastest and most capable small model yet.
That is the pattern across the whole industry this year. Often each release is cheaper and quicker than the last, and every lab is pushing the others. OpenAI’s GPT-6 Luna sits at exactly the same price and appears in Anthropic’s own comparison table. Competition is working, and we are the ones collecting the winnings.
Announcements are claims, so I ran it. Haiku 5.5 passed every hidden test in 70 runs of my coding suite. It cost about seven US cents a run.
The short version
I gave Anthropic’s new small model the same seven hard programming jobs I used on Sonnet 5.5, Opus and OpenAI’s GPT-6.1 Sol earlier this week. Each job is marked by tests the model cannot see.
Haiku 5.5 got every test right, on every run. So did Sonnet 5.5, 35 runs out of 35. The difference is the bill and the stopwatch. Haiku cost about a tenth as much. Sonnet was faster, by roughly 1.4 times.
My recommendation. For work where the instructions are clear and the tests are the referee, use Haiku 5.5. For anything where you sit and wait, or the problem is vague, stay on Sonnet 5.5. The best answer is probably both: a bigger model plans, Haiku does the typing. The section on planning and execution shows how to set that up in Claude Code. I tested the main routes headless.
One late entrant: OpenAI’s GPT-6 Luna, the same price class as Haiku, ran the suite for about two cents in under five minutes. It dropped tests in 2 of 14 runs. More on that below.
The detail follows, for anyone who wants to check my working.
The same numbers as a table, for anyone reading in a feed reader that drops images:
| Setup | Cost, seven tasks | Wall time | Hidden tests passed |
|---|---|---|---|
| Haiku 5.5 in Claude Code | US$0.07 | 5 min | 70 of 70 runs |
| Sonnet 5.5 in Claude Code | US$0.73 | 3 min | 35 of 35 runs |
| Opus 5.5 in Claude Code | US$1.33 | 4 min | 14 of 14 runs |
| GPT-6 Luna in Codex | US$0.02 | 4.3 min | 12 of 14 runs |
| GPT-6.1 Sol in Codex | US$0.38 | 7.5 min | 7 of 7 runs |
| GPT-6 Astra in Codex | US$1.77 | 11 min (before the speed-up) | 7 of 7 runs |
| Opus 4.6 in Claude Code | US$4.82 | 22 min | 12 of 14 runs |
This is the chart from my Sonnet against GPT-6.1 post with Haiku added. I moved cost to a log scale, otherwise Haiku sits on the axis.
Cost, speed and chattiness
Haiku 5.5 is priced at US$0.10 in and US$0.50 out per million tokens for prompts up to 100,000 tokens. Sonnet 5.5 is US$2 and US$10. That is a factor of 20 per token, and about 10 per suite, because Haiku uses more tokens to get there.
It writes about 65,000 output tokens a suite against Sonnet’s 22,000 to 28,000. It reasons out loud more and takes more turns on some tasks. On one task it needed about 80 seconds where Sonnet needed 28. Per-token price is not per-task price, so measure the task.
Speed is the soft spot, and also the noisiest number here. Summing the median time per task, Haiku took 299 seconds a suite and Sonnet 211. A same-day Sonnet rerun took 415 seconds because one task ran slowly. Call the gap 1.3 to 1.6 times and don’t build anything on the decimal.
One more trap. Claude Code 2.1.289 does not recognise Haiku 5.5 yet. It still runs, but its own cost figure came out at about US$1.40 to 1.50 a suite, around 20 times the real price. I recomputed from the token counts. If your usage screen looks wrong, that is why.
The other cheap model: GPT-6 Luna
Haiku’s real price rival is OpenAI’s GPT-6 Luna, listed at US$0.10 in, US$0.01 cached and US$0.50 out per million tokens (OpenRouter’s listing puts it level with Haiku). I ran it through the Codex CLI at medium effort, twice over the same seven tasks.
It was the cheapest of anything I have tested at US$0.02 a suite. It took 245 to 268 seconds, quicker than Haiku and slower than Sonnet, with about 15,000 to 17,000 output tokens a suite against Haiku’s 65,000. Codex also caches most of the prompt, so the input bill is tiny.
It was also the only cheap model to miss. In one run it failed a single hidden test on the H3 task (rejecting bad amounts). In another it failed 5 of the 11 tests on H6. Across both suites it passed 206 of 212 tests. That is 97%, which would be a fine score on most things and is the wrong score for work where the tests are the referee.
Two runs per cell is a small sample, so I would not call Luna worse than Haiku on this evidence, only less proven. The cost gap is real though: Haiku is about US$0.07 a suite and Luna about US$0.02. A planner and executor split, with Luna doing the typing and tests catching the slips, is the experiment I would run next.
What it does to a 5-hour window
On my Claude Pro plan, Sonnet 5.5 used six or seven points of the 5-hour meter per suite. I then ran eight Haiku suites in a row. The meter moved from 22% to 30%, so about one point a suite.
That is roughly 100 Haiku suites in a window against 15 to 18 for Sonnet, about six times more. Treat one point as a ceiling. The meter rounds to whole numbers and my own session draws on it too, so the true figure may be lower. Anthropic does not publish a per-model weighting, so this is a measurement, not a rule.
Plan with a big model, execute with Haiku
This is where the numbers get interesting. Planning is a small share of the tokens. Execution (editing files, running tests, fixing the typo it just made) is most of them. Put a strong model on the first and Haiku on the second and you keep most of the judgement and pay for very little of the typing.
Claude Code gives you three ways to do it. I tested each headless and read the usage report to see which model did the work.
The built-in route is opusplan. The opusplan model alias uses Opus in plan mode and switches to Sonnet to execute. Two environment variables change who does what. Point the Sonnet alias at Haiku and execution runs on Haiku. Point the Opus alias at Sonnet and planning runs on Sonnet:
export ANTHROPIC_DEFAULT_OPUS_MODEL=claude-sonnet-5-5 # optional: plan with Sonnet
export ANTHROPIC_DEFAULT_SONNET_MODEL=claude-haiku-5-5 # execute with Haiku
claude --model opusplan
In my tests the plan phase ran only on Opus (or only on Sonnet when remapped) and the execute phase ran only on Haiku. In an interactive session you press Shift+Tab to enter plan mode, approve the plan and Claude Code switches to the execution model. I could only test that handover headless, so check it in your own session. The plan lands as a markdown file under ~/.claude/plans. One catch: you have remapped the Sonnet alias for that whole session, so anything else that asks for “sonnet” gets Haiku.
The subagent route keeps your main session smart. Create .claude/agents/executor.md:
---
name: executor
description: Carries out an approved plan by editing files and running tests. Use proactively once a plan exists.
tools: Read, Edit, Write, Bash, Grep, Glob
model: claude-haiku-5-5
---
You execute plans exactly. Change only what the plan lists.
Run the checks. Report each step as done or blocked, and list any deviation.
Run your main session on Opus or Sonnet, plan in plan mode, then say “use the executor subagent to carry out the plan”. In my test the usage report showed Sonnet and Haiku side by side, and Haiku made the edits. A subagent starts with a fresh context, so it knows only what you put in the plan or the file it reads. That is a feature. Your main session stays clean for review.
Use the full model ID in the file. The haiku alias follows whatever your Haiku environment variable says, and the docs do not tell you which version that is. To push every subagent to Haiku at once, the docs show CLAUDE_CODE_SUBAGENT_MODEL=haiku with CLAUDE_CODE_SUBAGENT_MODEL_FORCE=1. I did not test that one.
The scripted route is two commands. Plan with claude -p "..." --model opus --permission-mode plan, then execute with claude -p "Carry out the plan in <file>" --model claude-haiku-5-5 --permission-mode acceptEdits. It is the most explicit and the easiest to log. It also loses the interactive review step, so I’d use it for batch jobs only.
Two things to know before you trust it
The executor will tidy things you did not ask it to tidy. In my toy test I asked it to add a flag to a one-line script. Haiku replaced the existing greeting with a new one. The main session noticed and reported it, which is exactly why you want the main model reviewing the diff. Write “change only what the plan lists, and report any deviation” into the executor’s instructions, and let your tests judge the result.
And I have not shown that planning helps. Haiku passed my suite on its own, with no planner at all, so there is nothing to improve. The suite cannot tell you whether a plan from Opus makes Haiku better on a vague, sprawling job. My guess, and it is a guess, is that it does, and that the saving lands on the execution tokens. Test it on your own work before you rely on it.
How solid is this
Anthropic’s own page is more modest than my results. It lists Sonnet 5.5 ahead on every benchmark, including Terminal-Bench 4.0 (70.6% against 39.2%) and says Sonnet and Opus remain the better choice for complex agentic coding. My suite is seven self-contained tasks with clear specifications, which is the easy end of that range. Read my result as “Haiku is enough for well-specified jobs”, not “Haiku equals Sonnet”. I also ran Haiku at its default setting. This is the first Haiku with an adjustable effort level, and I have not tried the other settings.
One suite, one harness (Claude Code for Anthropic, Codex for OpenAI). Each cell is 7 to 70 runs, and Luna’s is 14. Luna’s cost is tokens at list price, not a billed amount. The tests were written by Claude-family models. The speed gap is noisy and the quota figure is a ceiling. The orchestration tests used a toy task to confirm which model does what, not to measure quality.
What is solid is the pass rate. Over 100 runs between them, Haiku and Sonnet failed nothing. A model that costs a tenth as much and does not miss a test is hard to ignore, so try it on the boring half of your week first.
Sources:
- Introducing Claude Haiku 5.5 - Anthropic
- Pricing - Anthropic, Claude Platform docs
- Claude Haiku - Anthropic
- Model configuration - Anthropic, Claude Code docs
- Create custom subagents - Anthropic, Claude Code docs
- Permission modes - Anthropic, Claude Code docs
- Run Claude Code programmatically - Anthropic, Claude Code docs
- Manage costs effectively - Anthropic, Claude Code docs