Position: LLMs Can't Jump

A position paper arguing that large language models can’t make Einstein-style “leaps of intuition” has prompted debate over whether current AI systems are structurally limited to pattern-matching and deduction. Commenters question the claim that abductive, world-changing insights require sensory grounding or physical intuition, pointing to multimodal models, agentic systems, and recent mathematical results as evidence that capabilities may continue to scale. Many see the real open problem as how to rigorously define and measure such “jumps,” and whether future architectures or training regimes—not just bigger LLMs—will be needed to achieve them.

Scope of the claim “LLMs can’t jump”

  • Thread centers on whether current LLMs can perform “abductive” leaps (new foundational hypotheses) akin to general relativity’s key insights.
  • Many argue the paper is mostly a position piece, not backed by quantitative benchmarks or clear operationalization of “jumping.”
  • Others like the framing of induction vs deduction vs abduction, but question whether abduction is well-defined enough to justify “structural incapability” claims.

Sensory grounding, world models, and intuition

  • One line of argument: major theoretical leaps (e.g., relativity) depended on embodied, sensory-grounded thought experiments; text-only LLMs lack this.
  • Counterpoints:
    • Humans make abstract leaps in math and CS with minimal direct sensory link.
    • LLMs can run text-based thought experiments and already handle physics intuitions to some extent.
    • Multimodal models and agentic systems with tools or simulators may already approximate world models.
  • Several commenters think dedicated world models and physical feedback loops are promising, but current results are mixed or “abysmal” for general reasoning.

Historical and physics nitpicks

  • Multiple commenters argue the paper oversimplifies the history of relativity, Lorentz transformations, and Einstein’s influences.
  • Debate over how “intuitive” the equivalence principle really is; some find it obvious, others see it as profoundly counterintuitive.
  • Broader point: many great breakthroughs were incremental on top of an active research ecosystem, not singular mystical leaps.

Empirical tests and data/time limitations

  • Proposed experiments:
    • Train a modern LLM only on pre‑1980/1990 or Victorian-era text and see if it can “reinvent” LLMs or modern physics.
    • Use future LLMs (with older cutoffs) to reproduce recent “jump” papers.
  • Challenges noted: training data scarcity pre‑internet, leakage of modern concepts, and retrospective bias in what historical documents were preserved.

Capabilities, limits, and “goalpost shifting”

  • Some argue every past “LLMs can’t…” has become “LLMs can, poorly, then better with scale,” so blanket impossibility claims are suspect.
  • Others insist there are clear current limits: long-form math proofs, unsupervised codebase maintenance, reliable self-generated training data (model collapse), and genuine novelty in humor or science.
  • Several see LLMs as powerful “language calculators” that augment human intuition and creativity rather than replace them; humans+LLMs may jointly make the leaps even if models alone do not.