Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
Cognition’s new SWE-2 coding model, built by post-training on Kimi K3, is being marketed as rivaling top frontier models like GPT-Astra and Fable 5.1, but many engineers question whether its headline scores come from genuine capability or “benchmaxxing” against older, saturated benchmarks. Commenters point to its sharp drop on the newer Terminal Bench 4 and its closed weights, arguing that open models like DeepSeek V4.1 Flash offer better price–performance and more control. Others note that SWE-1.x was already quite usable, see value in Cognition running its own fine-tuned stack to reduce dependence on OpenAI/Anthropic, but are frustrated that SWE-2 is tightly tied to Cognition’s Devin tooling rather than standard APIs or open distribution.
Benchmarks, “Benchmaxxing,” and Generalization
- Many commenters question Cognition’s self-published FrontierCode benchmark, noting it’s internal, not reproducible, and positions SWE‑2 as “rivaling Astra/Fable.”
- The gap between Terminal Bench 2.1 (≈93%) and Terminal Bench 4 (≈27%) is repeatedly cited as a red flag for overfitting/“benchmaxxing” and poor generalization to new problems.
- Others point out that top models like Sol and GLM also drop sharply from TB2.1 to TB4, arguing this is an ecosystem-wide issue and that TB4 is not yet saturated.
- Some see coding benchmarks as increasingly low-signal marketing numbers; they advocate building custom workloads instead.
Model Origin, Training, and Capabilities
- SWE‑2 is described as “post‑trained from Kimi K3,” leading to discussion of post‑training vs. distillation vs. fine‑tuning, and how much extra capability RL/post‑training can add.
- There is debate over whether a ~2.8T-parameter model can realistically match ~10T-class models; some are skeptical RL alone can close that gap.
Access, Product Integration, and Lock‑In
- SWE‑2 is currently accessible mainly via Cognition’s Devin ecosystem (CLI, desktop/cloud), not through standard APIs or aggregators like OpenRouter.
- Some users dislike needing a bespoke CLI/harness just to test a single lab’s model and request a simple chat/completions-style API.
- There is confusion over pricing: claims that SWE‑2 is free for a month via CLI, discounted in cloud, and free on-device, but some users report being charged.
Open vs Closed Weights and Chinese Base Models
- Several commenters ask whether SWE‑2 is open‑weights; disappointment is voiced if it’s closed, especially when compared to DeepSeek V4.1 Flash and other open Chinese models.
- Discussion notes that many US startups are now post‑training or wrapping strong Chinese open models (Kimi, GLM, etc.) rather than training from scratch.
Devin/Windsurf Product Experience
- Experiences with Devin/Windsurf are mixed: some enterprise and individual users find it useful as a coding agent, especially with integrated frontier models.
- Others report severe UX and reliability issues with the CLI/desktop harness, lack of key features, and overhyped earlier demos, leading to distrust.
Economic and Competitive Context
- Commenters see SWE‑2 partly as a way for Cognition to reduce dependence on OpenAI/Anthropic APIs.
- There is broader debate about cost-performance tradeoffs, with many praising DeepSeek 4.1 Flash as a strong, cheap alternative and predicting pressure on closed labs’ pricing.