ARC-AGI Leaderboard

A new ARC-AGI leaderboard showing Anthropic’s Opus 5 far ahead of other models has triggered scrutiny of how reliable AI benchmarks still are. Commenters question whether models are being “benchmaxxed” through training on similar puzzles, hidden prompt scaffolds, or leaked test data, and debate whether ARC-AGI’s game-like tasks meaningfully measure general intelligence or just performance on a narrow genre. The thread also surfaces concerns about closed-model data retention promises, the exclusion of harnesses and agents from official scores, and a growing gap between impressive benchmark results and perceived day-to-day usefulness of frontier models.

Opus 5’s ARC-AGI-3 performance and credibility

  • Commenters note a huge gap between Opus 5 and other models on ARC-AGI-3; some find this “crazy” after trying the tasks.
  • Others question plausibility: current open attempts (e.g., Kaggle) reportedly get ~2% with heavy compute, and some “99%” claims on public subsets are viewed as unproven.
  • Several suspect targeted training on ARC-AGI-3–style tasks rather than a broad “general intelligence” jump.

Benchmaxxing, contamination, and private datasets

  • Persistent concern that benchmarks are easy to “benchmaxx” via memorization, genre-specific training, or hidden instructions.
  • Some argue Opus 5’s behavior (perfect on template-like tasks, worse on novel ones) looks like training on genre-specific data, not general reasoning gains.
  • Others point out the difficulty of proving contamination; private/semi-private splits exist but may still leak through traces, harness reading, or user-shared data.

Harnesses vs single-prompt evaluation

  • Debate over allowing “harnesses” (tool-using wrappers) around models.
  • One side: harnesses dominate performance, so benchmarks measure the wrapper, not the model; allowing ARC-specific harnesses breaks “G” in AGI.
  • Other side: humans use tools; forbidding harnesses makes the benchmark less relevant to real-world agent systems. Some suggest letting models build their own tools in-session.

Data retention, Fable, and trust

  • Fable is absent because its data retention policy reportedly prevented safe use of the semi-private set.
  • Multiple commenters distrust major labs’ “no training” guarantees, though one notes a zero-retention claim held up in court.
  • Some raise the difficulty of distinguishing “never stored” from “stored then deleted.”

User experience vs benchmark gains

  • A subset of users feel new models benchmark better but don’t feel substantially more useful than older ones, or even feel “worse” due to constraints and hedging.
  • Others report clear improvements (e.g., in niche technical reasoning) and attribute perceived stagnation to hedonic adaptation.

Is ARC-AGI a good benchmark?

  • Critics say ARC-AGI-3 over-indexes on game-like puzzles, assumes human gameplay conventions, and is poorly aligned with text-centric LLM strengths.
  • Supporters argue that game-like, novel interactive tasks are precisely what’s needed to test generalization and tool-free reasoning.

Costs, incentives, and AGI rhetoric

  • Some balk at reported evaluation costs (up to tens of thousands of dollars).
  • There is cynicism that massive financial stakes and “AGI/ASI” marketing create strong incentives to overfit or cheat on benchmarks.
  • A few argue AGI is still 10–20 years away and that current systems are advanced pattern-matching, not true general intelligence.