China's AI pricing war, receipts pending
Zhipu's viral $0.07 GLM-5.2 price already 10x'd, one reply claims
Zhipu's GLM-5.2 went viral at $0.07 per million tokens, a 95% cut posters called "almost free." One reply says the price rebounded 10x as a stunt for eyeballs. Practitioners add that GLM-5.2 never tested as frontier-grade.
Emmanuel Fabrice Omgbwa Yasse AI-assisted
2026-08-10 · 5 min read

On August 10, 2026, an X post chasing Chinese AI news put a number on the timeline: Zhipu's flagship GLM model, it claimed, had dropped 95% in price and now undercut DeepSeek. The post, originally in Chinese and machine-translated, shows 136,000 views. Jun Song, whose bio flags him as an ambassador for Alibaba's Qwen team, compressed it into a one-liner: "$0.07/m token. This is basically using frontier AI for free."
The replies did what replies do. "Insane price drop!" "That's almost free, really." Then a counter-claim landed, and the thread stopped being a price announcement and became a dispute.
In this piece:
- The 95% cut that went viral
- Ten times the price, and a stunt accusation
- Cached or computed: what $0.07 actually buys
- The frontier label vs. real workloads
- China's AI pricing spiral
The 95% cut that went viral
You need the surrounding market to see why this spread. DeepSeek V4 landed in mid-July with time-of-day pricing: costs double during business hours, up to $1.74 per million output tokens, while off-peak cache-hit workloads get the cheap end. Claude Sonnet 5 launched at $2 per million input tokens and $10 per million output, introductory rates that step up to $3 and $15 after August 31. Anthropic has run this play before: Opus 5 launched at half the price too. Against that, a Zhipu flagship quoted at $0.07 per million tokens is not a discount. It's a different species of price.
That's why the post moved. The translated claim said GLM-5.2 now sat below DeepSeek after a 95% cut, and Song's math made it quotable in one line. What the thread never shows: Zhipu's price page, an official announcement, an expiry date. It's claim and amplification; the receipt never appears.
Ten times the price, and a stunt accusation
The counter came from Matt Penny: "Prices have risen again to about 10x what is shown there. They did it just so people create posts like this and get more eyeballs. They play'd ya."
If Penny is right, the thread reads differently. A 95% cut that holds at $0.07 is a headline. A price that dips to $0.07 long enough to harvest the posts, then settles near $0.70, is a publicity event with a meter running. Other replies fed that reading. Manuel De Ceglie, who was told the API now undercut Zhipu's coding plan, replied: "Are you saying that I would pay less via API than with the coding plan? That's crazy, deleting it was a good idea." A poster named Darryl went for the joke: "they made enough money this year with those coding plan price hikes lol."
None of this verifies anything. Penny is a single account with no attached rate sheet, and his claim contradicts Song's screenshot directly. But the contradiction is the content: a price engineered to convert, and an accusation that the engineering is the point.
Cached or computed: what $0.07 actually buys
The most technical pushback came in all caps, verbatim: "these prices are based of cached or computed!!!!!" That question decides whether $0.07 was ever a real number.
Every major API sells cache hits at a fraction of full-token price, and DeepSeek's off-peak cache-hit reputation is built on that gap. A per-token rate with no cached/computed split is a teaser, not a price. Kyle Airey asked the other practical question, whether the same rate applies on Alibaba Cloud, where much Chinese model traffic is routed. The thread answers neither.
| Price in play | Quoted rate | Source and status |
|---|---|---|
| GLM-5.2 (viral claim) | $0.07 per million tokens | Jun Song post and Jason Lee's 95% cut claim; not verified |
| GLM-5.2 (counter-claim) | About 10x higher after rebound | Matt Penny reply; not verified |
| DeepSeek V4 | Up to $1.74 per million output tokens in business hours | seventnews July 2026 coverage of the V4 launch |
| Claude Sonnet 5 | $2 input / $10 output per million, intro | Anthropic, through Aug 31, then $3 / $15 |
The frontier label vs. real workloads
The price fight ran into a quality fight. Bartek Rutkowski ran GLM-5.2 against the workloads he handles daily with GPT and Opus models: "With all due respect, GLM-5.2 isn't nowhere near to frontier models. I jumped on the hype, tested it with my workloads that I deal daily with GPT/Opus models and very quickly discarded GLM-5.2 as not worth the money and time." The double negative is his. Sergio Suave landed softer: "GLM 5.2 is pretty far below frontier AI, but new pricing still pretty good."
What frontier means right now is checkable against published numbers. GPT-5.4 and GPT-5.2 Pro both sit at 99 on AIME 2025 and 97 on HMMT 2025, with Claude Opus 4.6 and Grok 4.1 in the same rankings. GLM-5.2 is not on that table. It's not the only big open release to miss the top tier; Kimi K3, the largest open model ever, still didn't beat the best. At $0.07 you buy something; the replies say it isn't the top of the leaderboard. The price per token misses the point if the tokens themselves are padded; a diagnostic benchmark found over half of LLM reasoning is froth.
China's AI pricing spiral
Step back and the thread fits a pattern. MiniMax's M3 is open-weight and priced to undercut Western APIs: its Ultra tier costs ¥469/month (about $65) for 5.5 billion tokens, roughly three times Claude Max's $200/month capacity at about a third of the price. The company's H3 follows the same logic; it prices 2K output below a third of mainstream. DeepSeek V4 Pro undercuts GPT-5.5 by around 9x on output price under an MIT license. Alibaba Cloud's Qwen models are cited as especially popular in Japan and South Korea for local-language strength, and Qwen3.8-Max's weights go public next week. In that market, a flagship at $0.07 is plausible. A flagship flashing $0.07 to generate the post is equally plausible.
TendiesOfWisdom added the standard objection: "The price is sending all your data to China." Meme or concern, it's part of why a cheap Chinese model travels so fast across an English-language timeline.
The thread closes where it opened, on a price that no one in it can confirm. One reply treats GLM-5.2 as near-free access to frontier work; another calls the price theater; a third says the model was never frontier to begin with. They agree on one thing: the price is the story, and the story does the selling. "They play'd ya," Penny wrote. If Zhipu's rate sheet says otherwise, nobody in this thread has linked it.
Get the tech essentials in 3 minutes every morning
One email, every weekday, with what actually matters in AI and tech.