Top AIs still fail IQ tests
Claims that top AI models still “fail” standard IQ tests prompt debate over whether such tests meaningfully measure machine intelligence at all. Commenters argue that large language models are powerful next‑token predictors with impressive but shallow generality, hampered by vision and reasoning limits, memorization-heavy “reasoning,” and difficulty with truly novel problems. Many conclude that IQ scores and labels like AGI are less important than real-world usefulness and the rapid pace of capability growth, which could make current benchmarks obsolete.
Relevance of IQ Tests for AI
- Many argue IQ tests are a poor way to evaluate LLMs: designed and validated for humans, especially within a narrow ability band, not machine systems with different architectures.
- Others note IQ (and the g factor) is still the strongest empirically validated proxy for human intelligence and life outcomes, so it’s at least an interesting benchmark.
- Disagreements center on whether IQ mainly measures pattern recognition vs broader skills like creativity, problem-solving, and “fluid intelligence.”
- Some see poor LLM IQ performance as more damning for the tests than for the models; others see it as clear evidence LLMs aren’t “intelligent.”
Nature of LLM “Intelligence”
- One view: LLMs are “just next-token predictors” / lossy compressed storage of web text, good at imitating but not thinking.
- Counterview: perfect next-token prediction would require deep understanding of underlying processes; current models already show some emergent reasoning.
- Several note that humans also “hallucinate” plausible nonsense when out of their depth, blurring the boundary between human and machine behavior.
- There is no agreed definition of intelligence, leading to essentially semantic fights over whether LLMs qualify or whether “AGI” has been reached.
Vision and Test Format Issues
- Many argue the specific IQ examples mostly test vision/encoding, not reasoning.
- Suggestions: present Raven-style problems as text, JSON, or ASCII art instead of images; some experiments show models then often solve them correctly, though stochastically.
- Failing image-based questions is compared to testing a blind person’s IQ with purely visual puzzles.
Capabilities, Limits, and Hallucinations
- LLMs can match or surpass humans on some textual IQ-like tasks and some math benchmarks, but struggle with precise arithmetic, generalization beyond training data, and robust abstract reasoning.
- Research cited suggests much “reasoning” is actually pattern memorization, though some genuine reasoning may exist.
AGI, Hype, and Future Trajectory
- Opinions range from “we already have AGI at average-human-mimic level” to “LLMs are overhyped, inch-deep pattern machines.”
- Some stress practical usefulness (“can it do work?”) over philosophical labels.
- Others emphasize rapid progress, parameter counts still far below brains, and expect significant near-term advances, especially in vision.