How GLM built its own inference infrastructure
Chinese lab Z.ai has built a large-scale inference stack for its GLM-5.3 models on more than 100,000 domestically made AI accelerators, prompting debate over the strategic impact of China achieving competitive, non‑NVIDIA infrastructure. Commenters weigh the technical achievements and aggressive software optimization against user complaints about pricing, speed, and token limits, and contrast GLM’s open-weight posture with closed US models like Claude. The thread also delves into allegations of large‑scale distillation from Anthropic’s models, the effectiveness of US chip export controls, and whether this shift will erode Western AI hardware and model dominance over time.
Chinese AI accelerators & geopolitical implications
- Many see running GLM-5.3(-Flash) on >100k Chinese accelerators as a major strategic shift: China now has its own chips plus large-scale inference.
- View that “most users don’t need frontier models” — strong, cheap, slightly-behind models plus domestic chips could be an “asteroid event” for Western labs.
- Debate on whether Chinese electricity is cheaper; one commenter cites data claiming industrial power is actually pricier than in the US.
- Some argue US export controls backfired, forcing/accelerating domestic Chinese chip development and efficiency innovations; others say China was always going to do this.
Pricing, token plans & value
- Several users say Z.ai’s early coding plans were “comically subsidized” and then sharply repriced, now feeling too expensive or too limited in tokens.
- Others still find good value due to caching and high token allowances, especially on Max plans, but note you must heavily parallelize and/or use Flash to hit quotas.
- Confusion over how many tokens subscription tiers actually provide; some compare unfavorably to Claude, others argue per-dollar value is similar when mapped correctly.
Model quality, usage patterns & agents
- GLM-5.3 is frequently praised as top-tier among open weights, with good reasoning and relatively low hallucinations, though it can “lose the plot” without oversight.
- Heavy token users describe multi-agent setups (e.g., code generation, bug-finding, search/optimization pipelines) driving billions of tokens/month.
- Consensus that all current models need periodic human or higher-level model supervision; GLM is no exception.
Distillation, legality & ethics
- Lengthy argument over allegations that Chinese labs (including Z.ai peers) illegally routed traffic through a US model provider for massive-scale distillation.
- Disagreement on legality: some call it clear fraud and cyberattack; others say ToS violations aren’t necessarily illegal or see it as “stealing from thieves” given alleged copyright infringements by Western labs.
- Skepticism over how much such distillation can really extract from limited conversations; others note that the scale of attacks suggests labs believe it works.
Performance, infrastructure & optimizations
- Some users say Z.ai’s service feels “dog slow” despite the blog’s optimization claims; others attribute this to deliberate throughput tradeoffs and power-saving.
- Discussion of high cache hit rates (90–97%) for agentic workloads, making token allowances more generous than raw numbers suggest.
- Technical interest in GLM-5.3’s full 744B MoE version running locally by streaming experts from SSDs, albeit at very low token/s.
- Several note that US labs are also aggressively optimizing inference and even building custom chips; this is not unique to China.
Export controls, protectionism & long‑term impact
- Debate on whether export bans help US labs (more access to Nvidia) and hurt Chinese labs, or instead create a protected space for Chinese chip vendors to catch up.
- Some predict future competition from cheap Chinese GPUs will erode Nvidia’s margins; others stress China’s remaining lithography/EUV bottlenecks.
- Broader reflections on industrial policy, trade wars, and the inevitability of large emerging economies catching up technologically.