
For six days in August 2026, the most-used AI model on the internet had no name, no price tag, and no public lab behind it. It was simply called Ox Alpha. In that stretch it processed 62 trillion tokens across the OpenRouter and OpenCode platforms, and it did so running entirely on Chinese-made GPU chips from three companies the United States has placed on its export-restricted Entity List. Then, on August 26, 2026, Zhipu AI's international arm Z.ai confirmed what the tokenizer fingerprints had suggested: Ox Alpha was GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series, running on a network of 100,000 domestically manufactured chips.
This is the story of that milestone, what it proves about China's chip industry, and the large caveats that come with every number in it.
Key takeaways
- Zhipu AI's GLM-5.3-Flash processed 62 trillion tokens in a stealth public trial in August 2026, running on domestic Chinese chips, according to the company.
- The chips came from three US-sanctioned vendors, Huawei, Hygon, and Moore Threads, per Zhipu and media reports. These attributions are company-sourced, not independently audited.
- On OpenRouter alone, the model logged over 11 trillion tokens in its first 72 hours, briefly handling about 31 percent of the platform's weekly coding traffic.
- The business momentum is real: Cambricon, Hygon, and Moore Threads all posted triple-digit or near-triple-digit revenue growth in the first half of 2026.
- The caveats are equally real: no independent auditor has verified the hardware claims, China still depends on imported high-bandwidth memory, and inference volume is not the same as frontier-model parity.
What Actually Happened
The sequence, as reported by TechTimes, IndexBox, and Chinese financial media: an unnamed model appeared on OpenRouter, a marketplace that ranks AI models by usage, and on OpenCode, an agent-focused coding platform. Over roughly six days it consumed 62 trillion tokens. OpenRouter's own figures showed it leading all coding models on the platform, with 10.3 trillion tokens in a single week representing roughly 31 percent of weekly activity, the largest launch the platform had ever seen.
On August 26, 2026, Z.ai confirmed the model's identity as GLM-5.3-Flash and stated that the stealth trial ran exclusively on a network of 100,000 chips manufactured in China. The next day, Zhipu's stock climbed more than 12 percent to close at HK$1,160 in Hong Kong. One detail with outsized significance: Moore Threads, one of the three chip vendors involved, completed a Day-0 hardware adaptation for the model within a day of its public debut, suggesting a maturing software ecosystem around domestic silicon, not just working hardware.
Why does a token count matter more than a benchmark? Because benchmarks are run in labs and tokens are burned by real users. Sixty-two trillion tokens of measured inference traffic answers the question US export policy has been asking since 2022: can China's domestically produced chips sustain frontier-class AI inference at real-world scale? For at least one high-profile workload, the measured answer was yes.
Why 62 Trillion Tokens Matters
Context turns a big number into a meaningful one. Chinese AI models have now led US models in weekly token usage for 22 consecutive weeks. In the week of September 21 to 27, 2026, Chinese models processed 62.22 trillion tokens against 14.2 trillion for US models, according to OpenRouter estimates cited by Chinese financial press. Back in June, Chinese models in the top 20 processed 98 trillion tokens in a month versus 53 trillion for US models, an 85 percent lead.
At the national level, Chinese officials said in mid-2026 that average daily token usage had reached several hundred trillion, up from 100 billion at the start of 2024. JPMorgan has forecast Chinese AI inference token consumption could grow roughly 370-fold between 2025 and 2030. The direction is consistent across every source: China is making AI usage cheap and ubiquitous at a pace nothing in the West currently matches.
The strategic backdrop is straightforward. US export controls have progressively cut off China's access to Nvidia's best hardware. Nvidia disclosed H20 sales of $4.6 billion before new licensing requirements took effect in April 2025, then reported zero H20 sales to China-based customers the following quarter. Beijing's response has been to mandate domestic adoption: in May 2026, China extended its state-backed "secure and reliable" technology certification to AI training and inference chips for the first time, with Huawei Ascend, Cambricon, Moore Threads, Iluvatar CoreX, Biren, MetaX, Kunlunxin, and Hygon all making the list. That list functions as a procurement catalogue for government bodies and state-owned enterprises, guaranteeing domestic chipmakers the scale to fund their next generation.

The Money Behind the Milestone
The financial results suggest an industry hitting an inflection point, not just a publicity moment. Cambricon reported first-half 2026 revenue of 5.996 billion yuan, up 108.13 percent year over year. Hygon guided to first-half revenue of 8.5 to 9.3 billion yuan, up 55.56 to 70.20 percent, with net profit up 41.50 to 52.32 percent. Moore Threads posted first-half revenue of 1.736 billion yuan, up 147.42 percent and already above its full-year 2025 total, while its net loss narrowed 95.73 percent to just 11.56 million yuan, bringing it close to profitability. Moore Threads is also pursuing a Hong Kong listing after its STAR Market debut, where its shares rose 425 percent at listing and its market value topped 350 billion yuan at peak.
The buildout is accelerating in parallel. In September 2026, JD Cloud announced a 100,000-GPU commercial cluster built on Moore Threads chips, priced by GPU-hour for enterprise customers. In July 2026, Sugon unveiled its own 100,000-card domestic supercluster, the Sugon 8000 Dengfeng, running on Hygon chips. Two vendors, two 100,000-card systems, announced weeks apart: that is Beijing demonstrating, at scale, that more than one domestic supplier can stand in for Nvidia. TrendForce expects domestic solutions to capture nearly 90 percent of China's high-end AI chip market in 2026.

What the Milestone Does Not Prove
Intellectual honesty requires the other half of the ledger, and it is substantial.
Nothing here is independently audited. The 100,000-chip figure, the all-domestic claim, and Moore Threads' reported 95 percent scaling efficiency on its 100,000-GPU cluster are company statements. As TechTimes noted of the Moore Threads claim, no independent Western auditor has verified the hardware or its firmware. Treat every vendor number as a claim, not a fact.
Memory remains the choke point. AI chips are only as good as the high-bandwidth memory feeding them, and that still comes overwhelmingly from three companies: SK Hynix, Samsung, and Micron. Hygon's own first-half results showed negative operating cash flow of 427.5 million yuan, a sign of how tightly supply-constrained the ecosystem remains. Domestic GPUs plus imported memory is not full self-sufficiency.
Inference is not training. Serving tokens efficiently is a genuine achievement, but training frontier models at the largest scales still favors Nvidia-class hardware and the software ecosystem around it. A country can burn enormous inference volume on agents, video generation, and customer service while still trailing at the frontier of model training.
Volume is not leadership. China's token dominance reflects aggressive pricing, open-weight models, and massive domestic deployment. It does not by itself prove Chinese models have closed the capability gap with the best US systems. Tokens measure usage, not intelligence.
Practical next steps
- If you follow semiconductors, watch JD Cloud's 100,000-GPU Moore Threads cluster: sustained commercial performance there would validate more than any press release.
- Track the memory story. Any Chinese breakthrough in high-bandwidth memory would remove the biggest remaining dependency.
- For investors, note the gap between market valuations and verified performance. Enthusiasm is running ahead of audited evidence across these names.
- For everyone else, the practical takeaway is simpler: expect Chinese open-weight models to keep getting cheaper and more capable, which means better AI tools at lower prices globally, regardless of who wins the chip race.
The bottom line
Sixty-two trillion tokens is the most concrete data point yet in the debate over China's chip independence: measured, public, real-world inference at frontier scale on domestic silicon. The revenue growth, the 100,000-card clusters, and the procurement mandates all point the same direction. But the claims remain company-sourced, the memory bottleneck is real, and token volume is not frontier parity. China has proven it can serve AI at staggering scale without Nvidia. Proving it can train the next generation without Nvidia is the test still to come.
Sources
- TechTimes: Sanctioned Chinese chips just served 62 trillion AI tokens at frontier scale (August 2026). https://www.techtimes.com/articles/325872/20260828/sanctioned-chinese-chips-just-served-62-trillion-ai-tokens-frontier-scale.htm
- IndexBox: GLM-5.3-Flash, Zhipu AI's new open-weight model runs on domestic chips (August 2026). https://www.indexbox.io/blog/zhipu-ai-releases-glm-53-flash-powered-by-100000-domestic-chips/
- TechTimes: Moore Threads claims 95 percent scaling on 100,000 GPUs, no independent auditor has verified it (September 2026). https://www.techtimes.com/articles/327151/20260910/moore-threads-claims-95-scaling-100000-gpus-no-independent-auditor-has-verified-it.htm
- TrendForce: China AI chip makers make waves, Cambricon 1H26 net profit surges 123 percent, Moore Threads eyes HK listing (August 2026). https://www.trendforce.com/news/2026/08/11/news-china-ai-chip-maker-cambricons-1h26-net-profit-surges-123-moore-threads-eyes-hong-kong-listing/
- StartupFortune: JD Cloud picks Moore Threads chips for a 100,000-GPU supercomputer (September 2026). https://startupfortune.com/jd-cloud-picks-moore-threads-chips-for-a-100000-gpu-supercomputer/
- abit.ee: China adds AI chips to trusted technology list as US export curbs bite (May 2026). https://abit.ee/en/processors/china-ai-chips-huawei-ascend-xinchuang-cambricon-moore-threads-us-export-controls-import-substitutio-en
- BeInCrypto: China's AI models process 98 trillion tokens, 85 percent above US (July 2026). https://beincrypto.com/china-ai-models-overtake-us-token-use/
Quick answers
Frequently asked questions
01
What does 62 trillion tokens mean?
Tokens are the basic units AI models process, roughly fractions of words. In August 2026, Zhipu AI's model GLM-5.3-Flash, tested under the stealth name Ox Alpha, processed 62 trillion tokens on the OpenRouter and OpenCode platforms while running entirely on Chinese-made chips. It is a measured volume of real inference traffic, not a benchmark score, which is why analysts treat it as a meaningful data point.
02
Which Chinese companies made the chips?
According to Zhipu AI and media reports, the stealth trial ran on chips from three companies on the US export-restricted Entity List: Huawei (Ascend), Hygon, and Moore Threads. These attributions are company-sourced and media-reported, not independently audited. Other domestic players include Cambricon, Biren, MetaX, Iluvatar CoreX, and Kunlunxin.
03
Can Chinese AI chips replace Nvidia now?
Not fully. The 62-trillion-token trial proves domestic chips can sustain large-scale AI inference, but training frontier models still favors Nvidia-class hardware, and China remains dependent on imported high-bandwidth memory from SK Hynix, Samsung, and Micron. Domestic chips are on track to capture a large share of China's own AI chip market in 2026, but the global frontier gap has not closed.
04
What is GLM-5.3-Flash?
GLM-5.3-Flash is an open-weight AI model from Chinese developer Zhipu AI, confirmed on August 26, 2026. It was tested publicly under the stealth name Ox Alpha before its official debut, and Zhipu says it is the first natively multimodal model in the GLM-5 series. During its trial it became the most-used model on OpenRouter, briefly handling about 31 percent of the platform's weekly coding traffic.
05
Why is China pushing domestic AI chips so hard?
US export controls since 2022 have progressively blocked China's access to Nvidia's most powerful GPUs, from the H100 to the China-specific H800, and by 2026 Nvidia reported zero H20 sales to China-based customers in a quarter. Beijing has responded with procurement mandates, a 'secure and reliable' certification for domestic AI chips, and massive state-backed demand, making self-sufficiency an operational necessity rather than a slogan.



