GPT-5.6 Sol Pricing Cut by 50% on OpenRouter

OpenAI’s GPT‑5.6 Sol model is being offered at a temporary 50% discount via intermediaries like OpenRouter and Vercel, but not through OpenAI’s own API, prompting speculation about price wars, market segmentation, and data incentives. Commenters weigh Sol’s capabilities and quirks—especially its tendency to over‑engineer solutions—against rivals such as Anthropic’s Fable/Opus and cheaper Chinese models like Kimi and DeepSeek, with many saying quality is now “good enough” and price and limits dominate choices. The thread also highlights concerns over zero‑data‑retention guarantees, energy use from extreme token consumption, and the sustainability of current business models as inference pricing drops toward raw compute cost.

Scope of the Price Cut

  • 50% discount applies to GPT‑5.6‑Sol only via intermediaries (OpenRouter, Vercel AI Gateway), and apparently only for a limited time.
  • Native OpenAI API docs still show full pricing; Azure/other cloud routes are not discounted.
  • Discount appears limited to the standard “OpenAI” route on OpenRouter (non‑ZDR), not to ZDR or BYOK configurations.

Motives and Market Dynamics

  • Many see this as a market‑share play and early salvo in a token price war, especially against Chinese/open‑weight models (Kimi K3, DeepSeek, Qwen, etc.).
  • Some argue it’s about juicing usage metrics ahead of IPOs and impressing investors, not profitability.
  • Others frame it as standard price discrimination: charge “captive” direct API customers more, discount where competition and price sensitivity are higher.

Model Quality: Sol vs Alternatives

  • A large group finds Sol 5.6 extremely capable: strong at debugging, low‑level hardware/driver work, reverse engineering, planning complex projects, and long‑horizon code reviews.
  • Another group sees 5.6 as a regression from older OpenAI models: overcomplicates simple tasks, hallucinates extra changes, “enterprise over‑engineering.”
  • Some say Grok, Kimi K3, DeepSeek v4 Flash, and Gemini 3.7 Flash are as good or better for their use cases at lower cost.
  • Others still rate Anthropic’s top model (esp. Fable) higher for architecture, long‑horizon orchestration, and following STYLE/AGENTS files.

Overengineering, Effort Levels, and Workflow

  • Common pattern: use Sol on higher “effort” levels for planning, architecture, and gnarly debugging; switch to cheaper/faster models (Luna, open weights, Mai‑Code) for implementation.
  • Many complain Sol and Fable both “do too much”; careful effort selection, terse instructions, and skills/agents files are seen as necessary hygiene.
  • Some report extremely heavy usage (hundreds of millions to billions of tokens/day) via agents and Ultra modes; others see this as wasteful or ethically dubious.

Pricing, Plans, and Subsidies

  • Strong consensus that subscriptions are heavily subsidized versus API, and that API pricing is “for corporations.”
  • Some say API prices are still above pure inference cost; others note independent inference providers prove API is not obviously subsidized.
  • Broad expectation of a “race to the bottom” where reasoning becomes a commodity priced close to compute, potentially threatening big labs’ business models.

Privacy, ZDR, and Providers

  • Zero Data Retention is a major factor for many; some only trust large labs or specific providers (e.g., DigitalOcean, DeepInfra, enclave‑based services).
  • Certain Anthropic models (e.g., Mythos/Fable) are noted as incompatible with ZDR agreements; enterprises explicitly disable them.
  • OpenRouter itself offers an opt‑in 1% discount for allowing training on conversation data.

Guardrails, Safety, and UX Frustrations

  • Multiple users report Anthropic models frequently refusing benign or technical tasks (bio/chem false positives, “advanced” math, mundane code), causing downgrades and subscription cancellations.
  • Others say they almost never hit refusals, suggesting country/config‑dependent behavior.
  • Sol is perceived as more permissive: will assist with gray‑area reverse engineering and security‑adjacent work where rivals refuse.

Benchmarks, Evaluations, and “Vibes”

  • Some reference public leaderboards, but many distrust them and emphasize in‑house evals and “price per successful task.”
  • Long threads debate “taste” (style, readability, design sensibility) vs raw capability; many concede these judgments are highly subjective.
  • Several note that model discussions are increasingly “vibes‑based,” similar to old programming‑language flamewars.

Environmental and Bubble Concerns

  • One commenter calculates massive energy use for billion‑token‑per‑day “tokenmaxxing” and calls it immoral; others counter with token caching, uncertain energy figures, or potential net‑benefit use cases.
  • Repeated worries that unsustainable capex plus falling prices could trigger an AI bubble pop, leading to GPU gluts and secondary‑market hardware dumps.