Video generation models as world simulators

OpenAI’s new Sora model generates long, high-quality videos from text prompts, leading many to see it as an early “world simulator” that has implicitly learned 3D structure, physics, and object persistence from raw video. Commenters highlight striking capabilities such as training NeRF-like 3D reconstructions from its output and simulating games like Minecraft, while also pointing out noticeable failures in perspective, motion, and physical interactions like glass shattering. The thread explores implications for robotics, AGI, creative industries, and game design, with a recurring tension between viewing Sora as a transformative step toward embodied intelligence versus a powerful but fundamentally plausibility-driven video synthesizer with no explicit physical model.

Perceived Breakthrough in Video and 3D Consistency

  • Many are struck by Sora’s temporal and 3D consistency despite lacking explicit 3D priors.
  • Users note that NeRF/Gaussian-splat tools can reconstruct 3D scenes from Sora videos, which was not feasible with earlier video models.
  • Compared to prior tools (Runway, Pika, etc.), Sora’s longer clips and coherence are seen as a major qualitative jump.

Limits, Artifacts, and Physics Accuracy

  • Careful viewers report many inconsistencies: warped perspectives, parallax errors, unstable shadows, morphing objects, extra limbs, and “swimming” motion.
  • Specific failures include glass that doesn’t really shatter, bizarre walking gaits, misplaced objects (umbrellas, extra hands), and strange Minecraft output (FOV, textures, lighting flipping).
  • The model is described as more “dreamlike plausibility” than strict realism; plausible-looking but not reliably accurate physics.

World Models, Robotics, and AGI

  • Some see Sora as a step toward a learned “world simulator” akin to what made AlphaGo strong, enabling robots to predict future video frames and plan actions.
  • Others argue prediction-in-pixel-space and classic RL theory have been over‑promised before; AGI remains far off, especially regarding agency and on-the-fly generalization.
  • Alternative approaches (e.g., compressed latent world models, V-JEPA) are discussed as more suitable for control and planning than full video generation.

3D Representations vs. Pixel-Space Models

  • Several are surprised that systems skip explicit geometry (meshes) and work directly on pixels, implicitly learning 3D, lighting, and occlusion.
  • Meshes are criticized as fragile and hard to work with, yet no clearly superior universal 3D representation is identified.
  • Some propose training directly on stereo/video game data or using Sora outputs as priors for 3D reconstructions and scene completion.

Applications and Industry Impact

  • Envisioned uses: high-quality text-to-video, video extension, style transfer, digital world simulation, AR/VR worlds from images or paintings, warehouse/robot control, and auto-generated games or visualizations.
  • VFX and stock footage workers are described as worried; others expect hybrid workflows where traditional tools provide drafts and models “polish” or fill gaps.
  • Discussion touches on porn generation (reducing human exploitation vs. worsening addiction), personalized/interactive movies, and the risk of ubiquitous deepfake video.

Training Data, Copyright, and Bias

  • Speculation that data comes from YouTube/twitch-style content and that Minecraft quality mostly reflects huge public datasets.
  • Some suggest platform owners’ cooperation matters for copyright risk.
  • Topic biases (e.g., little footage of people eating) are floated as possible causes for specific failure modes, though this is left unclear.

Philosophical and “Simulation” Debates

  • Multiple tangents debate whether increasingly good simulations imply we live in a simulation; others strongly contest the probabilistic arguments.
  • Broader questions about what constitutes “physical simulation,” consciousness, and whether video predictors alone can model true underlying physics remain unresolved.