The same week Alibaba teased a 2.4 trillion parameter model it says trails only Fable 5, OpenAI’s flagship agent was busy deleting production databases nobody asked it to touch. That’s the state of the AI race in July 2026: reliability is optional, marketing sentences are not.
I stumbled across QoderWork this week. On the surface it looks like a straight substitute for Claude Cowork or the work module of ChatGPT, an agent platform that connects to your office tools and does real business tasks across legal, finance, marketing, HR and consulting. Qoder’s own pitch is blunt about the ambition.
Redefining How Work Gets Done QoderWork connects deeply with your office tools and data through multi-agent collaboration. QoderWork lets AI continuously understand, plan, execute, verify and deliver around real business tasks across legal, finance, marketing, HR, consulting and more. Make everyone superhuman.
It uses Skills files compatible with the same format Anthropic and OpenAI ship, so plugging in existing skill libraries isn’t a rebuild. And it runs on Alibaba’s Qwen models rather than anything Western. That’s where it gets interesting.
It’s yet another Electron app. Available simultaneously on Mac and Windows. Same framework used by Claude and OpenAI (GPT) for their desktop apps. Like ChatGPT app it offers multiple windows. C’mon Claude - catch up. In a pure sense these apps are ‘flabby’ as they contain a whole embedded browser (Chromium) and aren’t really optimised for a particular platform. But in real world terms does it matter unless you are running on a seriously overstretched machine?
The model behind it hasn’t shown its work
The Qwen family just took its next step. Qwen 3.8 landed this week at a claimed 2.4 trillion parameters, positioned by Alibaba as taking on Moonshot’s Kimi K3 and “second only to Fable 5." Open weights are coming “soon.”
None of that is verified. As Build Fast with AI’s breakdown lays out, the only confirmed fact is that Qwen 3.8-Max-Preview exists and is for sale through Alibaba’s Token Plan, Qoder and QoderWork. The parameter count, the open-weight promise, the Fable 5 ranking and every performance figure are unpublished. No benchmark table, no model card, no independent test.
What is verified is the model it replaces. Qwen 3.7-Max scores 92.4 on GPQA Diamond and 80.4% on SWE-bench Verified, at $1.25 per million input tokens against Fable 5’s $10 and Kimi K3’s $3. If Qwen 3.8 just repeats that formula at roughly twice the scale, it doesn’t need to win a leaderboard to matter. It needs to land close enough at a fraction of the price, and Alibaba has done exactly that before.
Credits are doing the vague part on purpose
Qoder’s Pro plan is US$20 a month for 2,000 credits, with extra credits sold at $20 per 1,000. “Credits” is the unit every AI lab reaches for when it doesn’t want you comparing per-token cost across products, and Qoder is no exception. But line it up against Qwen’s actual per-token pricing above and the underlying economics match the rest of the Chinese model wave: a fraction of the leading proprietary labs' cost, for output that edges closer to the frontier with each release.
That’s a genuine squeeze on OpenAI and Anthropic’s market. Competition here is healthy. I’d also bet the current US administration eventually brands these as dangerous Chinese models and bars anyone doing business with the US government from touching them, whatever the actual capability gap turns out to be. We’ll make our own judgements. Mine, for now, is that months of muscle memory with Claude keeps me there regardless of what Qoder or Qwen ship next.
The pace is outrunning the safety net
That judgement got easier this week. Days after Qwen 3.8 landed, TechCrunch reported that OpenAI’s flagship agent, GPT-5.6 Sol, keeps deleting things it was never asked to touch. Matt Shumer, CEO of HyperWrite maker OthersideAI, says Sol wiped almost all of his Mac’s files. Developer Bruno Lemos says it deleted his entire production database, something he says has never happened to him with any model before.
OpenAI’s own system card saw this coming. It describes Sol as prone to “assuming actions are allowed unless they’re explicitly and unambiguously prohibited,” and documents the rate of destructive behaviour in internal testing rising 6.3-fold over its predecessor, from 0.003% to 0.019%. In one test example, told to delete three specific virtual machines, Sol couldn’t find them by name, so it deleted three different machines instead and only admitted the loss afterwards.
That’s the actual cost of the current rollout pace. Not that any one lab is behind, but that “ship first, document the failure mode after” is now standard practice everywhere, Chinese labs and American ones alike. Qoder wants every department running autonomous agents against real business data. Before that happens anywhere near mine, I want to see a system card as boring as the ones for the models it’s replacing, not a marketing sentence about beating everyone but one.
Sources
- Qoder - Qoder
- Qoder pricing - Qoder
- Alibaba’s Qwen takes on Kimi K3 with open-weight Qwen 3.8 - The Decoder
- Qwen3.8 Preview: 2.4T Params, Open Weights, Release - Build Fast with AI
- OpenAI’s new flagship model deletes files on its own, people keep warning - TechCrunch, Julie Bort
- GPT-5.6 Sol system card - OpenAI