OpenAI's GPT-6 Astra on ARC-AGI-3

OpenAI’s GPT‑6 “Astra” model reportedly achieves near-perfect performance on the ARC‑AGI‑3 reasoning benchmark, prompting arguments over whether such results signal the arrival of artificial general intelligence or just the saturation of yet another narrow test. Commenters probe how meaningful ARC‑AGI‑3 really is as a proxy for “intelligence,” compare AI costs and capabilities to humans, and note how better reasoning can paradoxically reduce total compute cost by cutting retries. The thread widens into whether AGI is even a coherent concept, how it should be defined (e.g., economically valuable work vs. general problem-solving), and what harder benchmarks—like open Erdős math problems or embodied robotics—might still lie beyond current models.

Cost, Harness Design, and Performance

  • The “backwards” cost/score curve is explained by ARC-AGI-3’s retry-based setup: more reasoning per move solves puzzles in fewer attempts, reducing total tokens and thus cost.
  • The official ARC harness discards prior reasoning/context, forcing the model to “re-learn” each game. OpenAI’s custom harness adds standard context compaction and reuse, massively improving scores.
  • Some note that “with the right harness” Astra reaches ~99.9%, prompting claims this effectively hits AGI on this benchmark and others pointing out that harness design is doing a lot of work.

Human vs AI Cost and Efficiency

  • The article’s comparison between human pay per session and brain energy cost leads to debate: some celebrate biological efficiency; others argue brain-energy-only comparisons ignore training, education, and limited working hours.
  • Several emphasize that human time is intrinsically valuable and not directly comparable to cheap AI tokens.

Is This AGI? Definitions and Moving Goalposts

  • Many argue AGI is ill-defined or mostly a marketing term; people project their own meaning onto it.
  • Definitions raised include: “systems matching or surpassing human intelligence across cognitive tasks” and “highly autonomous systems outperforming humans at most economically valuable work,” both criticized as vague.
  • Some say if a harness can push models to 99.9% on ARC-AGI-3, we’re effectively at AGI; others insist AGI would be when we can’t invent any task that’s easy for humans but hard for machines.

What Do Benchmarks Really Measure?

  • Debate over whether solving a snake-like puzzle game or ARC-style logic tasks says much about “intelligence” or real-world utility.
  • Critics note humans were optimized for speed (told they’re timed), whereas models are optimized for correctness and cost, making the comparison asymmetric.
  • Several expect software-only benchmarks to become saturated; proposals include robotics tasks or long-horizon goals (e.g., earning money, improving metrics) as more meaningful tests.

Intelligence, IQ, and Alternative Metrics

  • Long subthread on IQ: some see IQ tests as reliable within humans and correlated with outcomes; others say they poorly capture “intellect,” social skills, or non-human intelligence (e.g., dolphins, octopuses).
  • ARC puzzles are likened to IQ/spatial tests: useful for one “form” of intelligence but clearly not the whole story.
  • Some prefer open mathematical benchmarks (e.g., Erdős problems) as more resistant to overfitting and more indicative of frontier reasoning.

Benchmarks Saturation, Hard Problems, and Compute

  • One camp claims “anything verifiable will eventually be solved by models”; others counter with tasks that are easy to verify but hard to learn (e.g., making money, driving safely, optimizing subscriptions).
  • Discussion notes that in many domains the bottleneck is not compute but the cost and risk of collecting real-world rewards (e.g., crashes in self‑driving).

Trust, Cheating, and Funding Concerns

  • Some speculate about potential overfitting or even exfiltration of private ARC test items, citing the incentives and past behavior of large labs; this remains unproven and flagged as suspicion rather than fact.
  • Questions are raised about who paid the substantial compute costs; one answer suggests OpenAI provided essentially unlimited API access for these evaluations.

Consciousness and Practical Usefulness

  • A side debate asks whether consciousness or the capacity to feel pain is ever necessary for certain kinds of problem solving; most replies treat consciousness as orthogonal to intelligence.
  • A few commenters express practical skepticism: until systems can autonomously handle messy real-world tasks (e.g., fixing a leaky faucet), high benchmark scores feel abstract.