Grok 4.6
Grok 4.6, xAI’s latest large language model, is being positioned as a near–state-of-the-art competitor to Anthropic’s Fable and OpenAI’s GPT‑5.6, with users praising its speed, concise style, and cost-effectiveness, especially in coding and security workflows. Many remain skeptical of benchmark claims and note that real-world performance, particularly on complex reasoning and long-tail tasks, still lags top models in some areas. A major undercurrent centers on trust and ethics: some welcome Grok as a third frontier lab that pressures prices and guardrails, while others refuse to adopt it at all due to concerns about xAI’s safety policies, past CSAM/deepfake scandals, and Elon Musk’s political influence.
Model performance & benchmarks
- Many see Grok 4.6 as a substantial upgrade over 4.5; some claim “Fable‑like” or near‑frontier performance, especially on coding and reasoning benchmarks.
- Others report it still trails Fable and often Opus 5 in “judgment” and long‑tail tasks; several say benchmarks overstate real‑world capability (“benchmaxxing”).
- Comparisons vary: some put Grok 4.5 between Claude Sonnet and Opus, others say it codes at Opus‑4.8 level; several note Opus 5 feels worse than 4.8 despite better scores.
- There’s debate whether any model actually matches Fable; some users running their own code-review benchmarks find Fable clearly ahead.
Cost, limits & efficiency
- Grok is praised as fast and cheap, especially via Cursor, with generous usage relative to Anthropic.
- API pricing vs Sol/Qwen is contested: Sol often wins on token efficiency and speed, though Grok is closing the gap; some note pricing pages lag behind the 4.6 announcement.
- A few users say SuperGrok usage seems to be consumed faster with 4.6 than with 4.5.
Real‑world usage reports
- Coding: mixed but generally positive. Some use Sol for planning + Grok for implementation; others now consider Grok 4.6 good enough for both.
- Several report Grok finding real security issues and being more willing to test exploits locally (within constraints) than competitors.
- Others find 4.6’s plans “rambling” or self‑contradictory and say 4.5 felt more coherent.
- Voice mode: multiple reports of a recent regression—much terser, less nuanced, worse at multi‑part questions, though latency improved.
Safety, guardrails & scandals
- A leaked default system prompt shows explicit rules against criminal help, exploits, and sexual content involving minors, plus instructions not to reveal these rules.
- Many criticize prompt‑only safety as weak; others note providers layer additional filters and classifiers.
- Past use of Grok for CSAM and sexual deepfakes is heavily discussed; company representatives say it’s against policy and being enforced, while critics argue enforcement was late or inadequate and cite ongoing lawsuits.
Trust, politics & adoption
- A significant group refuses to use Grok or pay xAI due to the owner’s politics, moderation choices, and controversies (e.g., “Mechahitler”, antisemitic/racist outputs, deepfake scandal, foreign‑policy impact).
- Some orgs reportedly ban Grok over data‑privacy concerns, despite using Chinese models via other hosts.
- Others argue criticism is exaggerated or partisan and emphasize that Grok is technically strong and offers less restrictive security discussions.
Compute, data & competition
- Thread repeatedly attributes recent “Fable‑level” clustering across labs to: massive new compute coming online, shared RL task vendors, distillation (sometimes from Fable/Mythos), and continuous internal training pipelines.
- Some argue there’s “no moat”: with enough GPUs, data, and standard techniques, any major lab can reach frontier quality; compute is seen as the main moat.
- SpaceX/xAI’s rapid datacenter build‑out and the Cursor acquisition (coding traces as post‑training data) are cited as key to Grok’s jump.
- Many welcome Grok as a third strong frontier lab alongside OpenAI and Anthropic; Google/Gemini is seen as lagging, while Chinese closed and open‑weight models (DeepSeek, Kimi, Qwen) are viewed as strong low‑cost competitors.
Interface & UX
- Grok Build CLI/TUI and speech‑to‑text get praise (fast, polished, “no yapping,” concise answers).
- Some prefer Claude/GPT artifacts and writing style; others prefer Grok’s blunt, less anthropomorphic tone.