Karpathy’s Pelican
A viral demo of an AI model generating a 3D Lord of the Rings scene from text prompts is being used to debate what makes a meaningful benchmark for large language models. Commenters contrast this new “one-shot” 3D animation with earlier SVG “pelican on a bicycle” tests, arguing over whether such stunts show real progress in reasoning, spatial understanding, and software creation or just produce ever more polished but low‑value “slop.” Many also question the economic, creative, and environmental costs of burning large token budgets for flashy demos, and whether AI systems are anywhere close to reliably building things that are useful, fun, or robust in real-world settings.
New benchmark: LotR 3D animation vs Pelican SVG
- Thread centers on replacing the “pelican on a bicycle” SVG test with a harder benchmark: generating a three.js animation of the opening of Lord of the Rings.
- Some find this a better stress test: long, spatial, narrative, requires code, physics, and camera work—not just a static drawing.
- Others argue LotR is a bad choice because models are saturated with its text and imagery, so they may be copying the movies rather than “understanding” the passage.
Value and design of benchmarks
- Supporters: new tasks should start with terrible performance and leave headroom; pelican SVGs have converged and are less discriminative.
- Critics: pelican isn’t “solved” (bikes still non-functional), so moving on is premature and driven by demo appeal.
- Several call for benchmarks tied to real-world tasks (auditing, UIs, backends, vending machines), not arbitrary stunts.
Quality, creativity, and “AI slop” in games/media
- Many note that AI-generated games and 3D scenes look impressive in clips but are boring, broken, or unplayable; likened to “demo porn” or a “dancing bear”.
- Skeptics doubt LLMs can meaningfully optimize for “fun” or creative game feel.
- Some see clear utility as a shader/graphics/code assistant, not as an autonomous game designer.
Hyper‑personalized entertainment vs shared culture
- One camp envisions “1:1” AI movies/games from high-level prompts (choose‑your‑own‑adventure at scale).
- Another pushes back: most people lack specific desires for custom content and value communal experiences (e.g., discussing the same franchise together).
Capabilities and weaknesses
- Evidence of strengths: generating non-trivial three.js/WebGL scenes, Blender integrations, procedural animations, text→music experiments.
- Persistent weaknesses: spatial reasoning (bicycles, pinball tables, 3D prints), real-world constraints, and inability to robustly audit its own outputs.
- Some note copyright refusals are inconsistent; others worry about heavy reliance on iconic training data (e.g., LotR films).
Cost, externalities, and “~free” framing
- The “
free” narrative is challenged: the LotR demo alone cost ~1M tokens ($10) and rests on large, resource‑intensive data centers. - Environmental and local-infrastructure impacts of data centers are raised as a counterpoint to treating large runs as effectively free.
Software and throwaway tools
- Several foresee “throwaway software”: rapidly LLM‑generated one‑off tools instead of polished products.
- Others argue durable, maintained software is still economically superior; constant regeneration for complex systems is impractical and token‑expensive.
Trust, hype, and messaging
- Some commenters are impressed by rapid capability gains; others see overclaiming, corporate marketing, and shifting narratives about timelines.
- There is ongoing tension between excitement about possibilities and frustration with hype, slop, and shallow, attention‑optimized demos.