GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost
Claims that the Chinese GLM‑5.3 model outperforms leading OpenAI and Anthropic systems at a fraction of the cost have prompted scrutiny of how LLMs are benchmarked and compared. Commenters question the validity of the cited tests (small, saturated, often trivial tasks, AI‑generated write‑ups) and note that safety refusals, prompt harnesses, and enterprise procurement and compliance often matter more than raw scores. Some users report impressive real‑world performance from GLM‑5.3, especially on tasks like reverse engineering, but many argue that open or Chinese models still lag frontier US models in general capability and are complicated by geopolitical and licensing risks.
Benchmark credibility and methodology
- Many see the benchmark as saturated: ~10 models scoring >95% suggests the tasks are too easy or not discriminative.
- Coding evals are criticized as trivial (few, simple Python tasks; 28 tasks total), so differences between models may not be meaningful.
- Several point out inconsistencies with other leaderboards (e.g., ArtificialAnalysis, LMSys Arena), and results like Haiku scoring very high or GPT 5.5 > 5.6 are seen as red flags.
- Some note that Fable’s low score is largely due to refusals, not raw capability; others argue that counting refusals as failures is fair for “real-world” evaluation.
- Overall, many commenters regard the benchmark and write-up as “slop” and not a serious basis for model choice.
GLM-5.3 capabilities and cost
- Multiple users report GLM-5.3 (and related Kimi models) as strong performers, particularly for reverse engineering, tool use, planning, and coding in harnesses.
- Some say it can replace higher-end closed models for their personal/dev workloads at significantly lower cost; others find it slower, overengineered, or not clearly better than GLM-5.1.
- A separate evaluation is cited where GLM-5.3 ranks around #6 and is not Pareto-optimal on price/performance versus Grok and Sol.
Open weights, safety, and refusals
- GLM-5.3 weights are not yet released; delay is attributed to safety nerfs, which some worry may reduce capability but others say can be fine-tuned away.
- There is broad frustration with US models’ safety filters and refusals, especially for security, reverse engineering, and “autonomous” coding tasks.
Enterprise and geopolitical considerations
- Several argue that enterprises will continue paying for US big-tech models due to integrations, procurement, and compliance, regardless of small quality gaps.
- Others highlight supply-chain and political risks in both directions: US government can restrict US models, and could similarly target Chinese models, affecting contractors.
- Non-US organizations may perceive higher risk with US providers.
Trust, AI-written content, and propaganda concerns
- Many believe the article and site are heavily or fully AI-generated and poorly edited; this strongly undermines trust in the results.
- Some call this “AI slop” and worry about AI-generated content flooding HN.
- A few mention possible astroturfing or propaganda to promote Chinese models, though this is not substantiated beyond tone and results.