Stable-Audio-Demo
AI-generated music and sound from Stability AI’s Stable Audio demo impresses many with its audio quality and rhythmic coherence, but listeners consistently note an “uncanny valley” feel, odd instrument constraints, and repetitive or directionless musical structure. Commenters see near-term value for background tracks and game sound effects, yet say it remains far from human-level composition or editable, production-ready songs. The thread also highlights frustrations over closed weights and restrictive licensing, alongside broader arguments about copyright, training data, and whether rapid progress will plateau or soon match professional human output.
Overall audio quality & “uncanny valley”
- Many listeners say it “sounds like music” but feels off: odd progressions, lack of human-like phrasing, and strange instrument behavior.
- Guitar parts violate physical constraints of real playing (impossible chord shapes, odd tunings).
- Drum solo: good tempo but unrealistic timbres and artifacts (cymbals changing mid-hit, stick/brush confusion, background noise).
- Several note low-bitrate / “swishy” compression-like artifacts and constant reverb.
Genre-dependent performance
- EDM / rave-style tracks are rated as the strongest output, partly because the genre tolerates repetition and synthetic textures.
- Meditation/ambient works are considered “fine” as that genre is already texture- and randomness-heavy.
- Disco and more harmonically rich styles often produce bizarre or unidiomatic chord progressions.
Comparison to other music models
- Multiple commenters see this as better than prior SOTA (MusicGen, MusicLM) in coherence and prompt following.
- Others think tools like Suno produce more human-like songs with better long-term structure and chord/melody development.
- Stable Audio is often described as “good demo but still a glorified loop library.”
Control, inputs, and editing
- Text-only prompting is seen as too coarse for serious composition.
- Strong interest in: MIDI input, ControlNet-style control, img2img-like “audio2audio,” humming-to-track pipelines, and neural synth/sampler hybrids.
- Editing is highlighted as a core unsolved problem; people want in-painting–like tools for audio sections and stems.
Use cases: music vs sound effects
- Several musicians find the music boring, static, and lacking transitions or development.
- Sound effects impress some, especially for prototyping and indie games, though others find specific examples (e.g., footsteps) weak.
- Concern: commercial game use reportedly requires an “enterprise” license, limiting appeal for small devs.
Access, licensing, and IP
- No public weights; code for training/inference is released, but not datasets.
- StableAudio’s commercial product is said to be trained on licensed content via a library deal; some speculate that IP risk explains closed weights and stricter game licensing.
- Thread branches into a long, mixed-view debate over copyright, fair use, training on unlicensed data, open-source ML viability, and potential future regulation.
Browser / implementation notes
- Site warns against Safari; some report it “works fine,” others mention Safari struggling with many HTML audio tags.