Anthropic appears to be A/B testing reduced effort levels in Claude Code
Anthropic’s Claude Code IDE is drawing criticism from developers who say its newer Opus and Fable models are slower, more verbose, and sometimes seemingly “downgraded” behind the scenes, driving up token usage and costs. Many report better behavior at lower “effort” settings or with older models, and some are switching to competitors or open-weight models amid fears of enshitification and opaque billing incentives. An Anthropic engineer responds that recent changes are A/B tests of internal effort mappings rather than quality reductions, but users remain concerned about undisclosed routing, fluctuating performance, and being used as test subjects while paying.
Perceived effort-level changes and model routing
- Several users suspect Anthropic is A/B testing “effort” settings in Claude Code, dynamically downgrading effort or routing to cheaper/weaker models under load or based on user tier.
- Some report chat models feeling “lobotomized” at busy hours, while direct API usage feels stable.
- There’s debate over whether models can reliably report their own effort level; some argue it’s just system-prompt driven and thus knowable, others are unconvinced.
Model quality, regressions, and verbosity
- Many describe Opus 5 (and Fable) as dramatically more verbose, florid, and “LinkedIn-esque,” with long chains of thought for trivial tasks.
- Reports include extreme overkill (e.g., 40+ minutes of unnecessary analysis for a simple config edit) and worse math or accuracy at higher effort levels.
- A recurring theme: Opus 4.6 was beloved; 4.8 and 5 are widely perceived as regressions despite better benchmarks.
- Some users find Opus 5 acceptable or even strong on low/medium effort, especially for complex, multi-repo, multi-document synthesis.
Token usage, billing incentives, and limits
- Strong concern that higher effort defaults, verbose reasoning, and sub-agents “lighting tokens on fire” align with per-token revenue incentives.
- Others counter this is economically irrational given tight compute and competition; quality-per-token is what retains users.
- Frustration with opaque, variable token costs and lack of clear, fixed resource controls; some analogize it to a vendor controlling your “gas pedal.”
Comparisons with alternatives and local/open models
- Multiple commenters report switching or partially migrating to Codex, DeepSeek, GLM, Qwen, or Chinese open-weight models, citing better speed, stability, or cost.
- Debate over value of $20 subscriptions vs investing in local hardware; some argue local is costly, others emphasize predictability and control.
Trust, transparency, and safety programs
- Complaints about being effectively “test subjects” without opt-out when Anthropic tweaks configs.
- Grievances about revoked cyber-verification access and silent downgrades or rerouting (e.g., Fable → Opus) without always being clearly surfaced.
- Widespread worry about “enshitification”: optimizing for benchmarks, guards, or margins in ways that degrade day-to-day usability.