Stable Audio Open 1.0: 47-second limit and the Community License

Stable Audio Open 1.0 makes up to 47 seconds of 44.1 kHz stereo audio, trained on Freesound and FMA. Read the license note, limits, and a Sume alternative.

4 min readSume
All posts

Stable Audio Open 1.0 generates at most 47 seconds of 44.1 kHz stereo audio, and its model card puts it under the "Stability AI Community License", with a pointer to stability.ai/license for commercial deployment. It cannot generate realistic vocals and is "better at sound effects than music". That makes it a fit for short effects and loops, not for a full song under a voiceover.

What are the numbers on the model card?

The card lists the specs below. The training set matters for licensing because every file was under a Creative Commons style license.

Stable Audio Open 1.0 model card, https://huggingface.co/stabilityai/stable-audio-open-1.0, read 2026-10-03.
ItemModel card says
Maximum duration47 seconds
Sample rate and channels44.1 kHz, stereo
Training files486,492 total
From Freesound472,618
From Free Music Archive13,874
Training licensesCC0, CC BY or CC Sampling+
LicenseStability AI Community License

What does it say it is for?

The intended use is research into the limits of generative models, plus audio exploration by practitioners and artists. The card lists limitations: no realistic vocals, English-language training that hurts non-English prompts, uneven performance across music styles and cultures, and some prompt engineering. It also says the model should not be used to create audio that makes "hostile or alienating environments".

How does that compare with a hosted music job?

A Music Router job has no duration field; length is steered in the prompt, for example "a 2-minute track" or section markers like [0:00-0:30] Intro: .... The Music 1.0 page says it runs on Google Lyria 3.5 and makes full-length structured songs up to a few minutes, at a fixed $0.125 per accepted generation. Routable ids are sume/music-auto, lyria-3.5 and lyria-3-pro.

The two are different tools. The open model is a file on your GPU with a 47-second ceiling, and you own the setup. The hosted job is an API call that returns a Sume-hosted audio URL.

When should you pick which?

Choose the open weights when you want short effects generated locally and are fine with the Community License. Choose the hosted route when you need a bed longer than 47 seconds, or when the next step is a Sume Timeline render, because the Timeline soundtrack field accepts Sume-hosted audio only.

  • Under 47 seconds, effects or loops, own GPU: Stable Audio Open.
  • A 90-second bed under a voice: a Music Router job.
  • Anything commercial on the open model: read the Stability license page the card points to first.

What can you do with a 47-second ceiling in a video?

Plenty of useful audio is shorter than 47 seconds: a transition whoosh, a product sting, an intro, a loop that you repeat. A longer bed can be built from loops. If you go that way, Timeline 1.0 has soundtrack.loop, but only for a Sume-hosted file, so a locally generated loop must be looped in your own editor.

The card also says the model is better at effects than music, which matches the 47-second shape. Treat it as a sound-design tool first.

Why does the training data matter?

The card says all 486,492 training files were licensed under CC0, CC BY or CC Sampling+, and that 472,618 came from Freesound. Open training data makes the origin easier to explain to a client, which is part of what buyers ask when they check a music tool. It does not replace reading the model's own license for how you may use the output.

Compare that with a hosted engine, where you read the provider's and Sume's terms rather than a dataset description.

Sources

Related posts

More in Models

All Models posts

Written by Sume