Let me clear a huge misunderstanding

Claims that OpenAI’s new Sora video model “understands” the physical world are prompting pushback from researchers who argue it mostly stitches together plausible-looking frames without robust causal or object permanence. Commenters contrast Sora’s latent diffusion approach with alternative representation-learning ideas like JEPA, debate what “understanding” and “world models” should mean for AI, and highlight current failures such as inconsistent physics, disappearing objects, and limited temporal coherence. Some see Sora as a powerful but fundamentally statistical generator hyped by marketing, while others think its approximations of reflections, motion, and behavior already hint at emerging, if imperfect, world modeling.

Term “representation” and theory confusion

  • Several comments note the paper uses “representation” in the ML sense (feature/latent representation), not in the strict algebraic “representation theory” sense from group theory.
  • One poster warns that conflating these leads to overclaiming about breakthroughs; representation learning itself is not new.

JEPA / V-JEPA vs diffusion and generative video

  • A key theme: contrasting JEPA-style predictive representation models with diffusion-based generative video models like Sora.
  • Supporters of JEPA argue:
    • Learning compact, task-agnostic abstractions from video is more efficient and useful than directly generating pixels.
    • Generating many plausible-looking sequences is different from learning causal, action-conditioned dynamics.
  • Others counter:
    • Diffusion models already operate in a latent space, so they also learn abstractions.
    • It is “unclear” from the thread whether JEPA’s representations are fundamentally different or just another flavor of latent space.

Does generative video “understand” physics and the world?

  • One side claims Sora-like models only produce plausible “moving pictures” without real understanding of objects, causality, or physics; pointing to:
    • Object disappearance or morphing when occluded.
    • Violations of basic constraints (people walking through fences, dogs through shutters, odd snow dynamics, time running oddly).
  • Others reply:
    • Long, coherent clips with reflections, consistent cameras, and actions suggest at least partial internal structure about the physical world.
    • Even approximate, “illusionary” understanding can be practically useful.
    • Future scaling and memory mechanisms may reduce current coherence failures.

Meaning of “understanding”

  • Extended debate over whether models “understand” concepts or only associate patterns:
    • One camp: models are just deterministic pattern compressors over datasets; calling that “understanding” is unjustified without a precise, non-anthropocentric definition.
    • Another camp: if a system can reliably use a concept (e.g., cats, wolves) to generate consistent language or visuals, that is a form of understanding for practical purposes.
    • Several note the term is undefined and contested; “unclear” how to measure understanding quantitatively.

Hype, criticism, and research branding

  • Some see the tweet and JEPA promotion as a sales pitch and as reactive to recent video demos, with muddled arguments about what’s “pointless.”
  • Others defend critical voices as needed counterweight to aggressive commercial hype around AGI and “world simulators.”
  • There is disagreement over the value of labels like “self-supervised learning” and “JEPA,” with some calling them useful organizing concepts and others dismissing them as rebranding of older unsupervised/matrix factorization ideas.

Microblogging and access issues

  • Multiple posters complain that Twitter/X now often requires login, hindering discussion and archival.
  • Some advocate linking archive snapshots or pasting tweet text; others criticize microblogging as structurally bad for nuance.