Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights
China’s Z.ai has confirmed that its Ox Alpha model is a new GLM-series variant whose weights will be released, positioning it as an open contender to models like DeepSeek and mid-tier proprietary systems from Western labs. Commenters report strong performance on long coding and agentic tasks relative to its presumed size, but note unstable behavior such as “doom loops,” inconsistent benchmarks, slow inference, and possible heavy quantization. The release is seen as part of a broader strategic push by Chinese labs to open-weight high-performing models, intensifying competition and driving down costs for developers who can self-host.
Model identity & weight release
- Thread agrees Ox Alpha is a new GLM-series model from Z.ai; company reportedly confirmed and said weights would be released “tonight” (GMT+8, but timezone caused confusion).
- Some link it specifically to the GLM‑5.3/“flash” line; exact details of what will be released (full vs smaller variant, license terms) remain unclear.
Model size and architecture
- No official parameter count yet; community guesses range from ~18B active parameters (if like GLM‑5.3 flash) to 200–300B total.
- Several commenters stress that usefulness depends heavily on size vs performance: if small and strong, it’s a big deal; if closer to very large models, it’s just “good but not special.”
Performance & benchmarks
- Reported DeepSWE score around 63%; some say this is solid but below top frontier models.
- Public benchmarks show mixed results: underperforms some small GPT models on LiveBench, but other (unofficial) sites claim near‑Fable‑level; those unofficial sites and runs are widely viewed as untrustworthy or too small‑sample.
- Multiple comments highlight that most LLM benchmarks run on too few tasks to be statistically meaningful.
Hands-on usage reports
- Many used Ox Alpha for multi‑day coding and refactoring tasks; consensus: code quality and long‑horizon work are strong, often better than comparable “flash”/cheap models but behind top‑end Claude/GPT tiers.
- Good at rewriting, explanation, UI work, and well‑structured write‑ups; several describe it as “fun” or more readable than some Anthropic models.
- Others found it underwhelming, especially on complex bash pipelines or visual tasks.
Distillation and training debates
- Ongoing debate whether Ox Alpha is heavily distilled from other proprietary models; some see it as likely, others want stronger evidence.
- Linked research shows black‑box distillation can approach teacher performance, but doesn’t settle questions about reasoning/long‑horizon behavior.
Deployment, harnesses & doom loops
- Significant issues reported with slow inference, timeouts, and “doom loops” (repeating commands or tool calls).
- Some say this is common for GLM models and worsened by aggressive quantization or weak agent harnesses; others report stable long‑running sessions in better‑designed harnesses.
Ecosystem, licensing & geopolitics
- Many see open‑weight release as a competitive response to DeepSeek and part of a broader Chinese push for open models, possibly encouraged by state policy.
- Discussion about restrictive licenses vs truly open ones; some are fine with restrictions as business necessity.
- Observations of growing Chinese model adoption in startups, cheap local hosting via non‑NVIDIA hardware, and increasing brand confusion (Qwen, Kimi, GLM, Ox, etc.), though several deny there’s real confusion among technical users.
- Some note hype cycles and coordinated promotion around Chinese models; others view models as commoditized and expect users to keep switching to whatever is best/cheapest.