Published on [Permalink]
Reading time: 5 minutes

GPT-5.6 Luna cost drops 80%

Source: OpenAI, Advancing the price-performance frontier with GPT-5.6, using Artificial Analysis Intelligence Index v4.1 results.

OpenAI has cut the API price of GPT-5.6 Luna by 80 per cent, only three weeks after releasing it.

The new price is US$0.20 per million input tokens and US$1.20 per million output tokens, down from US$1 and US$6. Terra has also become 20 per cent cheaper. Sol is unchanged.

The chart above is the arresting part. Luna at maximum reasoning scores 51 on the Artificial Analysis Intelligence Index. Claude Opus 5 at low reasoning scores the same. Yet Luna appears to cost roughly five cents per benchmark task, against about 35 cents for Opus.

That is approximately the same measured intelligence for around one-eighth of the cost once Luna’s more precise post-cut figure is used.

Gemini 3.6 Flash is close to Luna’s score and costs roughly ten times as much per task. GLM-5.2 Max is also around the same intelligence level and several times dearer. At the lower end, Luna reaches an Intelligence Index score of about 33 for a little over one cent per task.

This is not merely cheap for a frontier model. It changes what kinds of jobs are economical to give a capable model in the first place.

The logarithmic axis softens the shock

There is a small but important label under the chart: log scale.

On a logarithmic horizontal axis, equal distances represent multiplication rather than equal numbers of cents. The distance from one cent to ten cents is the same as the distance from ten cents to one dollar.

That is a sensible way to fit models with very different costs onto one chart. It also makes the picture easier to misread. The gap between Luna at about five cents and Opus at about 35 cents looks like a moderate horizontal separation. In ordinary money it is a sevenfold difference. A linear axis would leave the Luna points pressed against the left edge while most competitors sat far to the right.

The logarithmic scale is not deceptive. If anything, it is necessary to show the detail among Luna’s reasoning levels. But it visually civilises a price gap that is commercially rather wild.

Does the claim survive an independent check?

There is a useful way to check OpenAI’s arithmetic without accepting its chart on faith.

Artificial Analysis tested GPT-5.6 Luna at its original API price. Its third-party evaluation gave Luna at maximum reasoning an Intelligence Index score of 51 and a cost of US$0.21 per task. Artificial Analysis also reports that Luna produced 188 output tokens per second, making it one of the faster reasoning models it had measured.

OpenAI has reduced Luna’s headline input and output token prices to one-fifth of their former levels. If the model uses the same number and mixture of tokens, and its cache rates scale with the base price, the independently measured US$0.21 task therefore falls to about US$0.042. That rounds to the four-to-five-cent point shown in OpenAI’s new chart.

So the most important part of the analysis is reproducible:

There are qualifications. Artificial Analysis worked with OpenAI on pre-release evaluation of the GPT-5.6 family, although it publishes its benchmark methodology and runs the models through its own suite. Its Intelligence Index combines nine evaluations covering agentic work, coding, science, mathematics, knowledge and long-context reasoning. It is broad, but it is still a benchmark rather than a guarantee about any particular workload.

The cleanest independent test for a business would be smaller and more personal: take 50 to 100 real tasks already completed satisfactorily by the current model, run them through Luna with the same tools and quality checks, and record three things — pass rate, total tokens and human correction time. The last measure matters. A model that is seven times cheaper per API call is not cheaper if people spend twice as long repairing its work.

Still, the benchmark result is too large to dismiss as graphmanship. Even if a private evaluation found Luna somewhat less capable than the composite score suggests, OpenAI has created an enormous margin for error. At four or five cents per serious benchmark task, Luna does not have to be the best model. It only has to be good enough surprisingly often.

That may be the more consequential frontier.

Why massive price drop?

OpenAI claim that they’ve used GPT-5.6 Sol to optimise their infrastructure and reduce operating costs. This is ‘How’ they’ve cut costs. The ‘Why?’ comes down to the intense competition with increasingly capable open weight models from China as well as Microsoft’s own AI models which are bundled with their apps.

OpenAI are ‘going for the jugular’ in an aggressive price move which will immediately hurt Anthropic and throw the challenge down to Google.

Sources

✍️ Reply by email