Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
Claims that China’s open‑weight Kimi K3 model is competitive with Anthropic’s top‑tier Fable model have reignited debate over whether AI is becoming a commodity and how much benchmarks really say about real‑world coding and agent workloads. Commenters dissect Fireworks.ai’s “oracle routing” methodology, question potential marketing bias, and compare cost, speed, token efficiency, and guardrails across K3, Fable, and other frontier models. The thread also highlights broader tensions: open vs closed weights, US vs Chinese AI ecosystems, data‑privacy and safety trade‑offs, and whether increasingly capable but poorly controlled systems should be made widely accessible.
Overall performance / SoTA claims
- Many see Kimi K3 as near–frontier quality, roughly comparable to Fable/Mythos and earlier frontier models (Opus 4.8, GPT‑5.5).
- Others argue Fable/Sol remain clearly ahead, especially on harder math (e.g., FrontierMath Tier 4) and some complex coding.
- Several commenters say open Chinese models are “benchmaxxed” and lag on real‑world tasks and token efficiency; others report K3 outperforming Sol/Fable on specific coding problems.
- Arena.ai and other leaderboards sometimes contradict Fireworks’ ranking, driving skepticism.
Oracle routing and practical routing
- Fireworks’ “oracle routing” is criticized as a theoretical upper bound: they run tasks on both models, then retrospectively pick the cheapest correct answer.
- Commenters note this assumes you can reliably detect correctness, which is often as hard as the original task.
- Some want evaluation of real, predictive routers instead; others still see value in showing the optimization ceiling.
- For companies routing may save significant cost; for individuals per‑task routing may hurt cache use and consistency.
Cost, tokens, and efficiency
- Big tension between per‑token price and total tokens per task: K3 often uses far more tokens and turns than Fable, so per‑task cost and latency can swing either way.
- Prompt caching (high hit rates in multi‑turn agents) can make K3 much cheaper despite extra tokens.
- Some users burn through Kimi subscriptions quickly; others find it half the cost of Fable for similar work.
Open‑weight vs closed labs
- Strong enthusiasm for open‑weights: self‑hosting, less risk of model deprecation or “lobotomization,” and competitive pressure on US frontier labs.
- Some argue open weights actually help closed labs via distillation and architectural insights.
- Disagreement over whether Chinese models are mostly copying/distilling US models or innovating independently.
Privacy, governance, and jurisdictions
- Kimi’s platform TOS allows training on user content by default; no simple opt‑out like Claude, raising concerns for proprietary code.
- Some prefer zero‑data‑retention (ZDR) routes via aggregators; others want EU or non‑US hosting to avoid both US and Chinese regimes.
Safety, refusals, and “personality”
- Frontier models (especially Anthropic) are criticized for aggressive refusals on security, biology, and even routine backend code touching auth/scopes.
- Chinese/open models are praised for fewer refusals and more autonomy, but others worry this is the “race to the bottom” on safety.
- Many dislike “glazing,” anthropomorphic, or nanny‑like tones (e.g., models scolding users for profanity); some explicitly want terse, even “rude” assistant behavior.
Tooling, routers, and self‑hosting
- Popular gateways/routers mentioned: OpenRouter, Bifrost, LiteLLM, OmniRoute, oh‑my‑openagent, OpenCode, various coding harnesses (Claude Code, Kimi Code, ZCode, etc.).
- Self‑hosting full K3‑class models is seen as feasible only for large orgs with significant GPU clusters, though quantized variants lower requirements.
Geopolitics and markets
- Multiple comments frame Chinese open‑weight advances as eroding US AI dominance and potentially commoditizing frontier capabilities.
- Some see China’s openness as strategic undercutting of US hyperscalers; others emphasize it’s just “geeks doing geeky stuff.”
- Debate over why stock markets no longer react strongly: novelty has faded, and narratives around DeepSeek’s earlier impact are questioned.