$100 AI Music Video: Claude Fable 5 vs. GPT-5.6 Sol
AI agents were tasked with independently generating low-budget music videos for Bruno Mars’ “Uptown Funk,” stitching together clips from modern text-to-video models. Commenters find the results technically impressive but artistically poor: overly literal, incoherent, and deep in the uncanny valley, a vivid example of what many call “AI slop.” The thread wrestles with whether such tools will mostly flood culture with cheap, soulless media or, when guided by skilled humans, eventually become powerful assistants that lower production costs and enable new forms of visual storytelling.
Overall reception of the AI music videos
- Many viewers found all the videos “awful,” “cringeworthy,” and artistically empty; some said they’d rather have no video than these.
- A minority were impressed by the mere fact this is possible at all for ~$25–$100 and under an hour, and thought they were comparable to low-end human-made videos or “Robot Chicken”–tier content.
- Several noted they were occasionally “so bad they’re funny,” especially the absurd literal shots (watches, washing machines, retired dragons, etc.).
Literalism, storytelling, and “AI slop”
- Main artistic critique: videos are hyper-literal illustrations of each lyric, lacking narrative, theme, or reveal—more like charades or karaoke visuals.
- Commenters contrasted this with good music videos that:
- Tell a story loosely related to the song, not 1:1 literal.
- Use irony, parody, or cultural context.
- This literalism is seen as emblematic of LLMs in general: shallow, overconfident, and unable to infer deeper metaphor or “vibe” without heavy guidance.
Technical capabilities and limitations
- Visual fidelity is often photorealistic in stills, but motion is janky: bad physics, off-beat dancing, anatomical glitches, drifting characters, and uncanny human faces.
- Major limitation: lack of temporal coherence and consistent characters across shots; models struggle even over 8–10s clips.
- Sync to music is poor because the system only has timestamps and text, not true multimodal understanding of rhythm.
- Several note the orchestrator LLMs used older or weaker video models (e.g., Wan), while newer ones (Seedance, Kling, Chinese systems) reportedly do better.
Methodology and intended takeaway
- Thread emphasizes the experiment’s goal was testing “agentic” tool use: LLMs planning, prompting video models, and editing clips into a full video under a budget.
- Many argue this setup is intrinsically handicapped: the LLMs can’t truly watch video, iterate visually, or exercise aesthetic judgment.
Human-in-the-loop and better examples
- Multiple links to artist-guided AI videos show much better results when:
- Humans design the concept, curate generations, and edit manually.
- Models are constrained (e.g., stylized animation, robot characters, band’s own artwork).
- Consensus: AI video is currently usable as a tool for clips, effects, and B‑roll, but not as a fully autonomous director.
Impact on art, taste, and creative work
- Strong philosophical split:
- Some see AI as just another medium or tool (like Photoshop, drum machines, or autotune); generative art with human intent is still art.
- Others argue fully AI-driven content is “slop,” anti-human, and hollow because it lacks human struggle, context, and intent.
- Widespread fear that:
- Cheap AI slop will flood platforms, especially kids’ content and low-end ads.
- Middle-class creative jobs (music video directors, VFX, ad studios, session artists) will be squeezed, even if top-tier “auteurs” and big-budget projects remain.
- A few counter that mass media has long been dominated by formulaic “slop”; AI mainly lowers cost and changes who can produce it.
Future trajectory
- Some believe the current outputs are just an early, bad phase and will improve rapidly with better prompting, models, and budgets.
- Others think we’re entering an “Age of the Plateau,” where everything is technically impressive but aesthetically mediocre—and that flooding culture with auto-generated content may worsen overall artistic quality and discoverability.