Published on [Permalink]
Reading time: 7 minutes
Posted in:

Sonnet 5.5 is fast, ChatGPT is cheap

I ran the same seven hard coding tasks through five models. Every current model passed every test. The only things that differed were the bill and the stopwatch.

The short version

Think of it as a driving test for AI coding assistants. I gave five of them the same seven hard programming jobs and marked the results with tests they could not see.

Four passed every job: Claude Sonnet 5.5, Claude Opus 5.5, ChatGPT’s GPT-6.1 Sol and GPT-6 Astra. The older Claude Opus 4.6 got two of its 14 attempts wrong. So the choice is not about quality. It is about money and time.

Claude Sonnet 5.5 was the fastest by a distance, finishing all seven jobs in about three minutes. GPT-6.1 Sol took about 16. But Sol was the cheapest, at about 45 US cents for the lot against 73 cents for Sonnet. Opus 4.6 cost US$4.82 and took 22 minutes.

On the US$20 monthly plans, ChatGPT Plus let me do somewhat more before hitting the 5-hour limit, roughly 10% to 30% more than Claude Pro.

My recommendation. If you sit and wait while the AI works, choose Claude with Sonnet 5.5 for the speed. If you set a job going and walk away, or you want image and video generation as well, choose ChatGPT for the better value. If you are still on Claude Opus 4.6, as I was, switch to Sonnet 5.5. It costs about a quarter as much and is faster.

I run both. Two US$20 plans beat one US$100 plan for the way I work. If you code every day or earn a living from AI, the US$100 plan starts to make sense.

The chart at top summarises the results. Detail follows, for anyone who wants to check my working.

The same numbers as a table, for anyone reading in a feed reader that drops images:

Setup Cost, seven tasks Wall time Hidden tests passed
Sonnet 5.5 in Claude Code US$0.73 3 min 14 of 14 runs
Opus 5.5 in Claude Code US$1.33 4 min 14 of 14 runs
Opus 4.6 in Claude Code US$4.82 22 min 12 of 14 runs
GPT-6.1 Sol in Codex US$0.45 16 min 7 of 7 runs
GPT-6 Astra in Codex US$1.77 11 min 7 of 7 runs

The tasks were not toys. A weighted cache with expiry. A scheduler that picks the most valuable jobs for k machines. A cron engine that survives daylight saving. A regular expression engine written without the re module. A minimal text diff. Each was marked by hidden tests, with brute-force answers as the referee.

Claude models ran in Claude Code and the OpenAI models in Codex. Same prompts, fresh folders, no hand-holding.

I was wrong about sticking with 4.6

In July I argued for pinning the 4.6 generation. The new tokenizer meant more tokens for the same text, and I thought that would eat any saving.

The tokenizer point still stands. It just does not decide the bill. On the six tasks I measured token by token, Sonnet 5.5 used about a quarter of what Opus 4.6 did (0.64 million against 2.58 million), because it needs far fewer turns to get there. Opus 4.6 took 33 turns on one task where Sonnet took 4.

Across all 14 tasks (seven everyday ones in the style of my own session tracker plus the seven hard ones) Sonnet 5.5 cost about 28% of Opus 4.6. The gap ran from 58% on the everyday work to 15% on the hard suite. The harder the job, the more Opus 4.6 flailed.

Sonnet was also about seven times faster on the hard suite and failed nothing. Opus 4.6 slipped on two of its 14 runs. The one place it won was a code review, 23 seconds against 36, and I have not scored review quality blind, so I call that a draw.

So Sonnet 5.5 costs about a quarter as much, runs faster and beats Opus 4.6 on every measure I tracked. My July advice is out of date. Unpin it.

Opus 5.5 gave the same answers as Sonnet 5.5 for 1.8 times the money. On this suite I cannot find a reason to pay for it.

Sonnet 5.5 against GPT-6.1 Sol

On quality it is a dead heat. Sonnet passed 14 of 14 runs. Sol passed 7 of 7 in Codex and 6 of 6 in Claude Code.

On price and speed they split. Sol in Codex cost US$0.45 for the suite against Sonnet’s US$0.73, about 40% cheaper. Sonnet finished in 3 minutes. Sol took 16. Both numbers matter, and they matter to different people.

The odd result is the harness. The same Sol model cost four times more inside Claude Code (US$1.65 against US$0.39 for six tasks). It burned 2.4 million tokens instead of 0.5 million, ran more loops and spawned subagents nobody asked for. Pick the tool as carefully as the model.

GPT-6 Astra is five times Sol’s price per token and passed exactly the same tests. It cost US$1.77 for the suite. I would not pay that for coding.

How fast each US$20 plan fills a 5-hour window

Neither company tells you how big the window is, so I measured. I read each plan’s meter before and after one full suite.

One suite of Sonnet 5.5 took my Claude Pro session meter from 8% to 14%. A second pass took it from 14% to 21%. That is six or seven points a suite, or about 15 suites per window. The true figure is probably a point lower, because the Claude session driving the test draws on the same meter, so call it 15 to 18.

One suite of GPT-6.1 Sol took ChatGPT Plus from 0% to 5%. That is about 20 suites per window.

Weekly, both moved about one point per suite, which suggests roughly 100 suites a week each. The meters round to whole numbers, so treat that as a very wide range.

OpenAI at least publishes a range for Plus: 15 to 160 GPT-6.1 Sol messages per 5 hours. For current Pro plans I could not find an equivalent from Anthropic. The only official figures I found referred to Sonnet 4.

How solid is this

Not bulletproof. Each cell is one or two runs. My scheduled Claude tasks and the Claude session running the experiment also draw on the same meter. A control run of four minutes with no tests moved it by zero points, so that overhead is small, but it is not nil. It is one morning’s snapshot, and it covers Sonnet 5.5 and Sol only. Opus and Astra would drain a window faster.

My own first scoring run also had a bug. A test generator emitted an escape that Python reads as a bell character, so every model failed the same four tests. All of them were right and my referee was wrong. I fixed it and rescored.

Where each one wins

ChatGPT wins on value. It is cheaper per task by API pricing and gives somewhere between 10% and 30% more suites per 5-hour window. A Plus subscription also includes image and video generation. Anthropic’s models do not generate either, so if you want that, the choice is made for you.

Claude wins on speed, by a wide margin. If you sit and wait on the model, five times faster beats a 30% quota edge. If you hand over a batch and walk away, Codex is the better deal.

Why I run both

I get more from two US$20 plans than from one US$100 plan. When one meter turns red I move to the other. I wrote about the gap between Pro and Max in July and the arithmetic has not changed.

If I coded every day or relied on AI for commercial work, the US$100 plan would earn its place. I do not, so it does not.

Before you upgrade anything, run your own typical week through both meters for an afternoon. They will tell you more than any price list.


Sources:

✍️ Reply by email