Stable Audio Open 1.0: 47-second limit and the Community License
Stable Audio Open 1.0 makes up to 47 seconds of 44.1 kHz stereo audio, trained on Freesound and FMA. Read the license note, limits, and a Sume alternative.

Stable Audio Open 1.0 generates at most 47 seconds of 44.1 kHz stereo audio, and its model card puts it under the "Stability AI Community License", with a pointer to stability.ai/license for commercial deployment. It cannot generate realistic vocals and is "better at sound effects than music". That makes it a fit for short effects and loops, not for a full song under a voiceover.
What are the numbers on the model card?
The card lists the specs below. The training set matters for licensing because every file was under a Creative Commons style license.
| Item | Model card says |
|---|---|
| Maximum duration | 47 seconds |
| Sample rate and channels | 44.1 kHz, stereo |
| Training files | 486,492 total |
| From Freesound | 472,618 |
| From Free Music Archive | 13,874 |
| Training licenses | CC0, CC BY or CC Sampling+ |
| License | Stability AI Community License |
What does it say it is for?
The intended use is research into the limits of generative models, plus audio exploration by practitioners and artists. The card lists limitations: no realistic vocals, English-language training that hurts non-English prompts, uneven performance across music styles and cultures, and some prompt engineering. It also says the model should not be used to create audio that makes "hostile or alienating environments".
How does that compare with a hosted music job?
A Music Router job has no duration field; length is steered in the prompt, for example "a 2-minute track" or section markers like [0:00-0:30] Intro: .... The Music 1.0 page says it runs on Google Lyria 3.5 and makes full-length structured songs up to a few minutes, at a fixed $0.125 per accepted generation. Routable ids are sume/music-auto, lyria-3.5 and lyria-3-pro.
The two are different tools. The open model is a file on your GPU with a 47-second ceiling, and you own the setup. The hosted job is an API call that returns a Sume-hosted audio URL.
When should you pick which?
Choose the open weights when you want short effects generated locally and are fine with the Community License. Choose the hosted route when you need a bed longer than 47 seconds, or when the next step is a Sume Timeline render, because the Timeline soundtrack field accepts Sume-hosted audio only.
- Under 47 seconds, effects or loops, own GPU: Stable Audio Open.
- A 90-second bed under a voice: a Music Router job.
- Anything commercial on the open model: read the Stability license page the card points to first.
What can you do with a 47-second ceiling in a video?
Plenty of useful audio is shorter than 47 seconds: a transition whoosh, a product sting, an intro, a loop that you repeat. A longer bed can be built from loops. If you go that way, Timeline 1.0 has soundtrack.loop, but only for a Sume-hosted file, so a locally generated loop must be looped in your own editor.
The card also says the model is better at effects than music, which matches the 47-second shape. Treat it as a sound-design tool first.
Why does the training data matter?
The card says all 486,492 training files were licensed under CC0, CC BY or CC Sampling+, and that 472,618 came from Freesound. Open training data makes the origin easier to explain to a client, which is part of what buyers ask when they check a music tool. It does not replace reading the model's own license for how you may use the output.
Compare that with a hosted engine, where you read the provider's and Sume's terms rather than a dataset description.
Sources
Related posts
More in Models
- Sume auto or a pinned video model: what changes past 10 seconds
Auto video defaults to 720p and 8 seconds with 3 to 10 second clips. For 12, 15 or 30 seconds, 1080p or 21:9, pin a model id. How to decide on Sume.
- Is Sume's kling-3 Kling VIDEO 3.0 or Omni 3.0? Read the catalog row
Sume's catalog names kling-3 'Kling Video v3 Pro' and lists no reference inputs. Kling describes Omni 3.0 as reference-driven, so match by capability.
- Sume Voices page lists 16 languages, Sonic 44: which list applies?
The Sume Voices page offers 16 languages when you create a voice; Cartesia states 44 for Sonic. The TTS request language field is a BCP-47 code you set per job.
- Tamil, Telugu, Kannada, Malayalam TTS API: language codes on Sume
Sonic 3.6 lists bn, ta, te, kn, ml, mr, gu and pa beyond Hindi. How to send each on Sume TTS, what the Hinglish note does not cover, and the cost per script.
Written by Sume