Posts in: agents

Jev and the rise of the cheap decision model

A start-up built an AI that cannot write a single sentence, and within weeks OpenAI and Cloudflare had built one too. Here is what Jev does, why it costs a sliver of what ChatGPT-style models charge and who is already using it. All dollar figures are US dollars.

Not three weeks ago, almost nobody outside the developer world had heard of Jev. This week, TechCrunch reported that OpenAI had announced its own version, and the headline called it a Jev clone.

Continue reading →



GPT-6 Astra: cheaper coding, pricier thinking

The most interesting thing about GPT-6 Astra isn’t its intelligence score. It’s that it now lies half as often as its predecessor did.

OpenAI launched Astra today, and it’s rolling out to subscribers gradually over the coming week. Which means I can’t test it myself yet, and neither can you. For now the only read available is Artificial Analysis’s independent benchmarking, published alongside the announcement.

Continue reading →


The weekly limit cut Anthropic is calling an increase

My Claude experience keeps getting less satisfactory. Every week I bang up against the Cowork weekly limit. Now that limit is about to get smaller, not bigger, whatever the headline says.

The math Anthropic already admitted

On 29 August, Anthropic’s own developer account, @ClaudeDevs, posted this on X:

“Starting September 14, we’re permanently raising standard weekly limits in Claude Code by 25% for Pro, Max, Team, and seat-based Enterprise plans. Until then, the current 50% increase will be in place.”

Continue reading →


Siri AI closes the gap on Gemini, but not all of it

Siri AI just got tested against Gemini in the open, not on an Apple keynote stage. That is a very different test, and this time it held up.

Tech reviewer Stephen Robles ran the two assistants head to head this week, on real devices, doing real tasks: pulling calendar dates from a screenshot, digging through email for a specific detail, identifying a movie from a screen recording, ordering a coffee without touching the app. I wrote about the architecture behind the new Siri back in June, when it was still a demo. This is the first proper look at what it does with your actual phone.

▶ Watch on YouTube

The result is closer than most of us expected. Not a wipeout in either direction. A genuine contest, with each assistant winning on different ground.

Continue reading →



Token blending: why your next AI agent will use more than one brain

Nobody wants to be a model selector.

Most people do not want to decide whether a task deserves Fable, Sol, Opus, Sonnet or Haiku. They want the job done properly, quickly and without the invoice arriving as a small technical mystery.

That is where Token Blending comes in.

The idea is straightforward. Start an agentic task with the biggest, smartest and most expensive model. Let it understand the problem, make the difficult calls and create the plan. Then hand the defined work to a cheaper, faster model to execute.

Continue reading →


Token inflation: why I'm sticking with Claude 4.6

The same code file now costs 30% more to process than it did three months ago. The price per token didn’t change. The tokenizer did.

Anthropic’s post-4.6 models - Sonnet 5, Opus 4.8, Fable 5 - use a new tokenizer that cuts text into more pieces than the 4.6 generation did, and you pay per piece. A detailed analysis by Playcode, who measured every frontier tokenizer on identical files, found the new Anthropic tokenizer produces 1.36 to 1.73 times GPT’s token count on the same content. TypeScript is the worst case at 1.73x.

Continue reading →


Why Claude still ships as an Electron app

Claude Code can write native Swift good enough to ship a Mac App Store app. Anthropic’s own Claude desktop app is still built using Electron - a web browser (Chromium) wearing a costume.

Electron bundles a full Chromium browser and a Node.js runtime inside every app, so one codebase renders on Windows, macOS and Linux instead of three native builds. It’s why Slack, Discord and (until recently) Notion feel identical everywhere. It’s also why they’re memory-hungry, slow to open and never quite right: patchy keyboard shortcuts, no proper native menu behaviour, battery life your MacBook resents.

Continue reading →