
Chinese AI labs have stunned the world, again. In the space of about a month, Chinese AI start-up labs Z.ai and Moonshot have each launched a model that is nearly as intelligent as competitors from OpenAI and Anthropic but far cheaper. Silicon Valley start-ups are already using Z.ai’s GLM-5.2, and Moonshot’s Kimi-K3 likely isn’t far behind in that market.
The success of lower-cost Chinese AI models highlights how high prices are likely to put Western models at a major disadvantage in the coming era of agentic AI, which is more expensive to use than large language models (LLMs) by several orders of magnitude. A future dominated by AI agents is one where companies value the cheapest intelligence possible. Right now, it’s advantage China.
AI prices are calculated per token – a computing-resources unit that Nvidia CEO Jensen Huang has called ‘the new commodity’ due to its growing importance in AI economics. AI companies use tokens to meter AI usage, roughly measured by the number of words that users put into an AI model and the number of words it then generates. While ordinary internet users pay for a monthly subscription, developers, who are using agentic models to build the future of AI, are charged per token. Buying 1 million tokens for, say, ChatGPT, buys a developer access to that model for a certain number of tasks. The price of 1 million tokens, some for users’ input and others for the AI output, varies between models.
GLM-5.2 tokens are cheap: Z.ai charges US$1.92 (A$2.74) per million output tokens for GLM-5.2, whereas Anthropic charges $25 per million for its Opus 4.8.
AI agents dramatically compound these cost differences. These models don’t just generate content but can also perform tasks online. The number of tokens needed to use agentic AI is exponentially higher than for conventional chatbot-style LLMs. Recent estimates from a team from MIT, Stanford and Google DeepMind suggest that agentic tasks use up to 3,500 times as many tokens as a simple reasoning task. A single prompt and response for information on a normal LLM usually requires a couple of thousand tokens. A single agentic coding or research task often runs into the millions, with some agent users at the extreme end already using tens or hundreds of millions of tokens per day for tasks that take days to complete. The emergence of agentic AI is one of the reasons that demand for tokens is going through the roof. One major AI provider, OpenRouter, recorded that its customers burned through more tokens in one week in June than in an entire three-month period the previous year.
Cheapness is attracting institutions with high token usage. High token payments factored into Microsoft’s recent cancellation of its subscription to Claude Code and could well be a factor behind its sudden interest in using DeepSeek to power its agentic AI office helper, Copilot Cowork. AI agent pioneer Azeem Azhar, who augments his own research firm’s work with AI agents, has started using MiMo-2.5-Pro, a model from Chinese conglomerate Xiaomi. MiMo’s token costs are dramatically cheaper than Claude’s, making all the difference when Azhar is using more than 100 million tokens daily. ‘At the scale agents operate, even small cost differences compound into meaningful budget gaps,’ wrote Azhar on his Substack, Exponential View.
The international spotlight has focused mainly on the most intelligent AI models, with Silicon Valley moguls such as OpenAI CEO Sam Altman and Anthropic’s Dario Amodei aiming for super-intelligent AI. Amodei in 2024 hailed this as AI’s future, ‘a country of geniuses in a data center’ capable of dramatic scientific discoveries. But when choosing a model, a customer may be looking not for the smartest model but for something more practical, something that’s good enough for the task at hand and affordable. For example, a company looking for AI agents that can draft and reply to emails doesn’t need a genius in a datacentre. A few percentage points lower in intelligence may matter less than a few hundred dollars more in savings.
Western and Chinese AI labs are wrestling with the question of how to keep prices down even as their costs go up. Anthropic and OpenAI have heavily subsidised their user subscriptions, with sovereign wealth funds and venture capital firms footing the bill. The rise in token usage within China has created a corresponding rise in computing costs, as demand for AI services outstrips domestic supplies of AI chips. In March, state broadcaster CCTV reported that cloud providers including Alibaba Cloud, Huawei Cloud and Tencent Cloud had all raised their computing prices by around 30 percent in just 10 days. This came at a time when market competition was also encouraging China’s AI enterprises to cut prices and offer free service schemes to attract users, as it still is. In 2025, SiliconFlow, a Chinese third-party provider of AI models, spent 24 percent more on marketing than its entire revenue.
But even as both sides contend with the economics of AI, current prices from OpenAI and Anthropic will likely be unsustainable if the world moves into AI agents. Chinese AI companies are already well-positioned to undercut the US companies. Take, for example, DeepSeek, which announced a permanent 75 percent cut to token prices for its application programming interface. That means that for any number of tokens (in fact, a very large number), an agent driven by DeepSeek can cost as little as 1/34 as much as one using competitors from OpenAI or Anthropic. Research company Artificial Analysis this month made leading AI models perform 657 agentic AI tasks, meaning high token usage. These tasks were all based around office and administrative work, tasks that agents may be performing for companies in the not too distant future. The total cost across these tasks for Anthropic’s Opus 4.8 was nearly $1,000. For Z.ai, it was $270. Despite burning the highest number of tokens of any model tested (1 billion), DeepSeek-V4 Flash cost just $14 to do all these tasks.

There are still disadvantages to Chinese models. The Economist has reported that GLM-5.2 uses more tokens when performing complex tasks than a Western equivalent, meaning the savings aren’t always as big as the sticker price per 1 million tokens implies. Developers also take performance into consideration when choosing a model – that is, whether it can do a job well. Chinese models don’t always perform well: they can break when they are given demanding tasks that take several days to complete. DeepSeek-V4 Flash scored poorly on performance, completing fewer tasks correctly for Artificial Analysis than both Western and Chinese competitors.
However, this may now be changing. Released just last week, Kimi-K3 completed more of Artificial Analysis’s tasks correctly than any other model tested, at around a third of the cost of Anthropic’s Fable 5. Once a Chinese AI model manages to combine cheapness with good performance, Western AI will be in trouble.

Lower prices for Western AI are already on the horizon. Meta announced on July 9 that its new AI model, Muse Spark 1.1, would be available at $4.25 per 1 million output tokens. This is a step away from the company’s earlier free open-source plan but is still a price that undercuts those of Anthropic and OpenAI. OpenAI has launched several versions of its new model, GPT-5.6, including Luna, a ‘cost-efficient’ model that is cheaper to run agentically than Kimi-K3 and GLM-5.2. However, a cheaper model sacrifices some of its performance. Kimi-K3 scores nearly 10 percentage points higher than Luna for performance on Artificial Analysis’s agentic test.
Another solution to high token costs is to host an agent within your own computer hardware. Nvidia is already positioning itself to cut the costs of AI agents, releasing a new laptop that can run agentic workflows locally, without paying additional costs through an application programming interface. Prices have not yet been announced but are said to be around US$2,000, making them within reach of enterprises in the West.
But if the cost of running AI agents is a problem even for well-financed Western companies, it will be an insurmountable obstacle for many in the Global South. Governments of developing countries seek to harness AI for their own development but chafe against multiple resource constraints and costs. China’s leadership is targeting the Global South for deployment of its AI products. In a speech last week at the World AI Conference in Shanghai, President Xi Jinping said China would launch AI application cooperation centres within six regional intergovernmental organisations that together cover almost the whole Global South. An op-ed in the People’s Daily last year argued that the low cost and open weights of Chinese AI made the technology more accessible to developing countries in the face of ‘hegemonic’ Western AI. An expansion of Chinese AI products in the Global South allows the technology to become part of the foundations of the AI ecosystems of developing countries going forward, with all the increased geopolitical leverage that would entail.
If Western companies don’t want to be left behind by their cheaper Chinese counterparts, they must find a balance between intelligence and affordability.
Leave a comment