Published on [Permalink]
Reading time: 7 minutes
Posted in:

Sonnet writes better research drafts, Haiku writes them for a cent

TLDR

The test. Haiku 5.5 and Sonnet 5.5 each wrote nine research drafts from identical source packs, using one fixed template. Two AI judges compared them blind, 27 verdicts in all.

The result. Sonnet won 16 verdicts and Haiku 11. Sonnet is better at insight, accuracy and readability. Coverage is tied. Neither model invented a quote.

The twist. It depends on the sources. Sonnet won 15 of 18 verdicts on messy news-style packs by cross-checking and catching errors. Haiku won 8 of 9 on a dense documentation pack, where breadth is the job.

The price. About one US cent a draft for Haiku and nine cents for Sonnet.

My take. Haiku for first-pass drafts on well-organised sources. Sonnet when the sources disagree or might be wrong.

I handed the same pile of sources to Haiku 5.5 and Sonnet 5.5, eighteen times, and asked for a research draft. Sonnet’s drafts were better. Haiku’s cost about a ninth as much, and on one of the three topics it won eight verdicts out of nine.

This is the second half of the experiment I started in my coding post. Coding was the easy test, because hidden tests give a hard pass or fail. Research work has no such referee, so I had to build one.

What a research draft is

A blog post is meant to be read once and enjoyed. A research draft is the opposite: a document you file, come back to in a month and trust enough to build on. The model’s job is to bring the sources together, draw out insights and present it all the same way every time, so a reviewer can scan ten of them without relearning the layout.

So I wrote a fixed template. Six sections in order: Summary, Key findings (numbered, each tagged to a source like [S2]), Where the sources agree and disagree, Insights (each starting with the word “Inference” so you can tell opinion from evidence), Open questions and gaps, and a Source table with a reliability note for each. The rules were short. Use only the supplied sources, plain English, 600 to 900 words, Australian spelling.

The test

Three topics, each a pack of about five real web sources that I fetched and froze so both models saw identical material. One was Google’s launch of skills in Gemini and how it compares with other platforms. One was this week’s Haiku 5.5 launch and pricing, from news coverage. One was the Claude Code documentation on choosing models for planning and execution.

Each model wrote three drafts per topic, so 18 drafts in all. Every draft was then judged blind, side by side, by Claude Opus 5.5 (in both orders, to cancel the habit of favouring whichever draft comes first) and by OpenAI’s GPT-6.1 Sol (one order each). That is 27 verdicts. The judges scored each draft from 1 to 5 on fidelity to the sources, coverage, insight, template consistency, accessibility and traceability, then picked the one they would rather file.

The result

Sonnet won 16 verdicts and Haiku 11. Opus split 10 to 8 and Sol 6 to 3, so the two judges agreed on the direction. The averages over all 27 verdicts:

Dimension Haiku 5.5 Sonnet 5.5
Fidelity to sources 3.81 4.15
Coverage 4.19 4.19
Insight 3.89 4.19
Template consistency 4.07 3.89
Accessibility 4.04 4.33
Traceability 4.22 4.30

Coverage is a dead heat. Haiku is slightly better at following the template. Sonnet leads on fidelity, insight and readability, and those are the three that matter most for a document you plan to trust later.

The by-topic split is the interesting part. On the Gemini skills pack Sonnet won 8 of 9 verdicts. On the Haiku launch pack it won 7 of 9. On the Claude Code documentation pack Haiku won 8 of 9.

Why the topics came out differently

The judges' own reasons tell the story. On the two news-style packs, Sonnet earned its wins by cross-checking. In one verdict it caught a unit error in one source (a tenth of a cent where the figure should have been a hundredth) and flagged a stale Haiku 4.5 heading on another page. In another it noted that a source never confirmed what the other sources assumed. That is what a good research assistant does, and Haiku mostly did not do it.

On the documentation pack Haiku won because it covered more. It picked up auto mode support, the classifier model, the cost figures and the agent team multiplier, and it found a real tension between sources that led to a strong inference. A judge called Sonnet’s version more precise but thinner.

My guess, and it is a guess from three topics, is that Haiku does well when the sources are dense and well organised, so breadth is the job. Sonnet pulls ahead when the sources are messy, sometimes wrong and need to be argued with. That would fit the insight numbers. Sonnet averaged 3.7 insights a draft and Haiku 2.0, even though both wrote about 880 words.

Cost, speed and honesty

Haiku wrote its drafts in 66 seconds on average and Sonnet in 47. Haiku produced about 15,300 output tokens a draft against Sonnet’s 5,400, so it thinks out loud a lot more, the same habit I saw on coding. The bill still comes out far lower: about one US cent a draft for Haiku and about nine cents for Sonnet. Haiku’s figure is repriced from tokens, because Claude Code does not yet know the model.

I also checked whether either model made up its quotes. Neither did. Every quotation in Haiku’s 18 drafts appeared word for word in the sources, and so did Sonnet’s 12, once I discounted two false alarms from my own checker. No invented links in any draft. Haiku slipped three Oxford commas past the style rule in nine drafts and Sonnet slipped none.

Every one of the 18 drafts followed the template, with all sections in order. I flagged a handful of figures in each model’s drafts as not appearing in the pack, three for Sonnet and four for Haiku, but I have not checked those by hand.

What I would do

Use Haiku for the first pass on any pack of structured documentation, where the job is to gather and organise. It is cheap enough to run on everything. Use Sonnet when the sources are news, opinion or vendor claims that disagree with each other, because that is where its cross-checking paid off.

The combination I want to try next is Haiku for the draft and Sonnet for a short review that adds the insights and checks the numbers. That is the same planner and executor idea from the coding post, run in reverse. It is untested, so treat it as a prediction.

How solid is this

Three topics, three drafts each, and the two judges are themselves language models, one from each lab. The drafts on the same topic are not independent, so the real sample is closer to three than to 27. I wrote the template and the judging rubric myself, and a different rubric could shift the scores. The cost figure for Haiku is a recalculation. What I trust is the direction: Sonnet is the better analyst, Haiku is the cheaper clerk, and the right choice depends on how messy your sources are.

If you do this kind of work, run your own three topics before you pick. It costs a few dollars.


Sources:

✍️ Reply by email