The most interesting thing about GPT-6 Astra isn’t its intelligence score. It’s that it now lies half as often as its predecessor did.
OpenAI launched Astra today, and it’s rolling out to subscribers gradually over the coming week. Which means I can’t test it myself yet, and neither can you. For now the only read available is Artificial Analysis’s independent benchmarking, published alongside the announcement.
I’m leaning on it because it’s the best signal going, not because I’ve got any way to check their working myself. Trusting a third party’s numbers is the deal you make when a model launches faster than anyone outside the lab can poke at it.
The headline number is the price. GPT-6 Astra costs 2.5 times what GPT-5.6 Sol did, up from $4/$20 to $10/$50 per million input/output tokens. Cache reads still get the same 90 percent discount, cache writes still carry the usual 25 percent premium. The structure hasn’t changed. The base rate has.
What you get for that money splits into two different stories, depending on which Artificial Analysis index you look at.
A genuinely better coding agent
In the Coding Agent Index, Astra scores 67, roughly level with Claude Opus 5 and Claude Fable 5 running in Claude Code, and with Muse Spark 1.3 in Muse Code. Claude Fable 5.1 still leads outright at 70.
What’s new is how Astra gets there. It uses a third of the tokens GPT-5.6 Sol needed at max effort in the Codex harness, and a fifth of what Claude Opus 5 burns at its highest setting. At max effort it costs about the same as its predecessor while scoring two points higher, and per task it comes in at less than half the cost of Claude Fable 5 for an equivalent result.
That’s a real efficiency gain, not a rounding error. A model that does the same coding work in a third of the tokens changes the economics of running it inside an agent loop, where token counts compound fast across a long session.
A pricier, mixed step forward everywhere else
The Intelligence Index tells a different story. Astra scores 61, tied with GPT-5.6 Sol and with Grok 4.6, and five points behind Claude Fable 5.1. It even trails Meta’s newly released Muse Spark 1.3.
Astra does use about 10 percent fewer output tokens than its predecessor at max effort, setting a new efficiency frontier on that measure. But the 2.5x price rise swamps it: per task, Astra ends up 75 percent more expensive than GPT-5.6 Sol at the same effort level, for a model that scores the same.
The genuine win sits in a benchmark most people won’t have heard of. AA-Omniscience measures hallucination rate, and Astra’s dropped from 92 percent to 51 percent at max effort, alongside a 4 point gain in accuracy. That’s not a model getting cagier and refusing more questions to look safer. Accuracy went up at the same time as the lying went down. If I had to pick the one number in this release that means something, it’s that one.
Everything else is a mixed bag. AA-Briefcase, Artificial Analysis’s long-horizon knowledge work evaluation covering multi-week projects with thousands of source files, shows Astra gaining around 80 Elo points, with real improvement in rubric scores and analytical quality. Its Presentation Quality Elo went backwards though, GPT-5.6 Sol at max effort still leads there. Humanity’s Last Exam is up 6 points. GDPval-AA v2, which measures economically valuable tasks across 44 occupations, is down around 80 Elo points. Small regressions of 2 to 3 points turn up in the customer support benchmark τ³-Banking, the scientific coding benchmark SciCode and the long-context reasoning benchmark AA-LCR.
Worth flagging too: Artificial Analysis attached an asterisk to GPT-5.6 Sol’s coding score, noting they’ve observed a higher rate of content safety filtering on that endpoint since they ran the benchmark. It’s a small footnote, but it’s a useful reminder that even a benchmark run by a genuinely independent lab is a snapshot of one endpoint on one day, not a fixed law of nature.
What this actually means
Strip away the launch framing and Astra reads less like a smarter model and more like a specialised one. As a coding agent it’s a real upgrade: cheaper per task, dramatically more token efficient, competitive with the best in the field. As a general intelligence model it’s flat on the metric that matters most, and considerably more expensive to run there.
If you’re running agentic coding workloads, the token efficiency alone probably justifies the switch.
Lots of examples of 3D simulations online from early testers. Some amazing results but for me they are more gee-whizz than practical. Others have pointed out that Astra can stick as tasks for days on end, provided you can afford the token bill.
If you’re using GPT for everything else, the case is a lot thinner. OpenAI just charged 2.5 times more for a model that thinks the same and codes better, and that’s a trade worth reading the fine print on before you upgrade by default.
I’m pretty happy with GPT 5.6. It hasn’t let me down and seems quick and smart enough at the tasks I throw at it. Will be keeping a cautious eye on the ‘progress’ promised in GPT 6 Astra.
Sources:
- GPT-6 Astra benchmark thread - Artificial Analysis
- Artificial Analysis - independent AI model benchmarking