Grok 4.6

Grok 4.6, xAI’s latest large language model, is being positioned as a near–state-of-the-art competitor to Anthropic’s Fable and OpenAI’s GPT‑5.6, with users praising its speed, concise style, and cost-effectiveness, especially in coding and security workflows. Many remain skeptical of benchmark claims and note that real-world performance, particularly on complex reasoning and long-tail tasks, still lags top models in some areas. A major undercurrent centers on trust and ethics: some welcome Grok as a third frontier lab that pressures prices and guardrails, while others refuse to adopt it at all due to concerns about xAI’s safety policies, past CSAM/deepfake scandals, and Elon Musk’s political influence.

Model performance & benchmarks

  • Many see Grok 4.6 as a substantial upgrade over 4.5; some claim “Fable‑like” or near‑frontier performance, especially on coding and reasoning benchmarks.
  • Others report it still trails Fable and often Opus 5 in “judgment” and long‑tail tasks; several say benchmarks overstate real‑world capability (“benchmaxxing”).
  • Comparisons vary: some put Grok 4.5 between Claude Sonnet and Opus, others say it codes at Opus‑4.8 level; several note Opus 5 feels worse than 4.8 despite better scores.
  • There’s debate whether any model actually matches Fable; some users running their own code-review benchmarks find Fable clearly ahead.

Cost, limits & efficiency

  • Grok is praised as fast and cheap, especially via Cursor, with generous usage relative to Anthropic.
  • API pricing vs Sol/Qwen is contested: Sol often wins on token efficiency and speed, though Grok is closing the gap; some note pricing pages lag behind the 4.6 announcement.
  • A few users say SuperGrok usage seems to be consumed faster with 4.6 than with 4.5.

Real‑world usage reports

  • Coding: mixed but generally positive. Some use Sol for planning + Grok for implementation; others now consider Grok 4.6 good enough for both.
  • Several report Grok finding real security issues and being more willing to test exploits locally (within constraints) than competitors.
  • Others find 4.6’s plans “rambling” or self‑contradictory and say 4.5 felt more coherent.
  • Voice mode: multiple reports of a recent regression—much terser, less nuanced, worse at multi‑part questions, though latency improved.

Safety, guardrails & scandals

  • A leaked default system prompt shows explicit rules against criminal help, exploits, and sexual content involving minors, plus instructions not to reveal these rules.
  • Many criticize prompt‑only safety as weak; others note providers layer additional filters and classifiers.
  • Past use of Grok for CSAM and sexual deepfakes is heavily discussed; company representatives say it’s against policy and being enforced, while critics argue enforcement was late or inadequate and cite ongoing lawsuits.

Trust, politics & adoption

  • A significant group refuses to use Grok or pay xAI due to the owner’s politics, moderation choices, and controversies (e.g., “Mechahitler”, antisemitic/racist outputs, deepfake scandal, foreign‑policy impact).
  • Some orgs reportedly ban Grok over data‑privacy concerns, despite using Chinese models via other hosts.
  • Others argue criticism is exaggerated or partisan and emphasize that Grok is technically strong and offers less restrictive security discussions.

Compute, data & competition

  • Thread repeatedly attributes recent “Fable‑level” clustering across labs to: massive new compute coming online, shared RL task vendors, distillation (sometimes from Fable/Mythos), and continuous internal training pipelines.
  • Some argue there’s “no moat”: with enough GPUs, data, and standard techniques, any major lab can reach frontier quality; compute is seen as the main moat.
  • SpaceX/xAI’s rapid datacenter build‑out and the Cursor acquisition (coding traces as post‑training data) are cited as key to Grok’s jump.
  • Many welcome Grok as a third strong frontier lab alongside OpenAI and Anthropic; Google/Gemini is seen as lagging, while Chinese closed and open‑weight models (DeepSeek, Kimi, Qwen) are viewed as strong low‑cost competitors.

Interface & UX

  • Grok Build CLI/TUI and speech‑to‑text get praise (fast, polished, “no yapping,” concise answers).
  • Some prefer Claude/GPT artifacts and writing style; others prefer Grok’s blunt, less anthropomorphic tone.