Moonshot AI suspends new subscriptions due to Kimi K3 demand
Moonshot AI’s new Kimi K3 model has drawn such heavy demand that the company has paused new subscriptions to preserve service quality for existing users, prompting comparisons with OpenAI, Anthropic, and other major providers’ capacity strategies. Commenters report strong coding and long-context performance, especially in multi-agent and agentic workflows, but complain about slow responses and tight usage caps on lower-tier plans, which can burn through quotas quickly. The surge is also fueling broader arguments over open-weight models, pricing sustainability, infrastructure limits, and whether Chinese frontier models can erode the incumbents’ lead.
Demand spike and subscription pause
- Kimi K3 demand has pushed Moonshot AI near capacity; they paused new subscriptions to preserve quality for existing users.
- Many commenters praise this as customer-friendly compared with quietly nerfing limits, though some note they could also have restricted access to K3 or adjusted limits instead.
- Some existing subscribers report they can still upgrade or extend plans.
Pricing, quotas, and value
- Users on the ~$20/month tier report blowing through 5‑hour daily caps with a few complex tasks, especially with K3 and multi‑agent workflows.
- Confusion exists around multiple plan types (research vs coding, older vs newer naming), priorities, and overlapping caps (5‑hour, weekly, monthly).
- Several people feel lower tiers are poor value for K3 and recommend higher tiers or metered usage instead.
- Some think K3 is underpriced given demand; others compare it unfavorably to cheaper models like DeepSeek for some workloads.
Performance, speed, and workflows
- Many are impressed with K3’s quality, particularly for code review, PR review, and complex research reports (e.g., multi‑hour runs generating book‑length outputs).
- A major complaint is slowness under load, especially via official harnesses; some think the bottleneck is the harness or quotas more than the raw model.
- Users describe widely differing practices: some let agents run 30+ minutes or even hours unattended, others find long runs degrade quality and prefer short, tightly scoped sessions.
Model behavior, safety, and “slop”
- K3 is described as roughly Opus‑level capable for some coding tasks and less “sloppy” in phrasing, but opinions differ on whether it justifies higher costs versus competitors.
- Some say K3 is less censored than Anthropic’s models; others say it has become more censored than earlier Kimi versions and “sounds like Claude.”
Architecture and open‑weights implications
- Technically inclined commenters highlight K3’s heavy use of RNN/linear attention–style layers (delta attention), linking it conceptually to xLSTM, Mamba‑like, and Qwen‑style architectures.
- There is discussion that transformers were meant to replace non‑parallelizable RNNs, but new tricks allow parallelizable RNN variants.
- Open weights are seen as crucial: once released, third‑party inference providers and self‑hosting can mitigate capacity crunches and distribute compute.
Market and geopolitics
- Some speculate that US labs retain an edge mainly in scaling, reliability, and hyperscaler contracts, while Chinese open‑weights models pressure them on price and openness.
- There’s frustration that Europe hasn’t funded large‑scale xLSTM‑style projects or dedicated AI supercomputing, with detailed suggestions for EU‑backed chips and infrastructure; feasibility is debated.