Stable Audio 3 max length: six minutes, and how to reach it on Sume

Stability says Stable Audio 3 Medium and Large pass six minutes. Sume's Lyria music runs up to a few minutes; join separate takes for more with Timeline audio.

4 min readSume
All posts

Stability AI says Stable Audio 3.0 Medium and Large produce audio longer than six minutes, and Small up to two minutes. Sume's music runs on Lyria 3.5, which its docs describe as songs of up to a few minutes, with no duration field. To get six minutes from Sume you generate separate tracks and join them, which gives you several pieces in sequence, not one six-minute composition.

The Stable Audio facts are from Stability's announcement, read 2026-09-29. The Sume facts are from Music 1.0 and Timeline audio.

How long can each side go?

Lengths as stated by each source, read 2026-09-29.
SourceStated length
Stable Audio 3.0 SmallUp to 2 minutes (Stability)
Stable Audio 3.0 Medium and LargeMore than 6 minutes (Stability)
Sume Music 1.0 (Lyria 3.5)Up to a few minutes, steered by the prompt (Sume docs)

How do I set the length on Sume?

Put it in the prompt. The docs suggest a phrase such as "a 2-minute track" or section markers like [0:00-0:30] Intro: .... duration and duration_seconds are rejected. Treat the length as a target and check the audio you get back.

How do I make something longer than one track?

Timeline audio joins Sume-hosted audio with operation: concat. It takes 1 to 20 ordered parts, joins them in the sample domain with no silence at the seams, and can produce audio up to 1800 seconds. It costs $0.01 flat per job. Import any outside file first with POST /v1/media-imports; parts must share a channel layout.

curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: long-score-001" \
  -d '{
    "operation": "concat",
    "parts": [
      { "url": "https://media.sume.com/artifacts/artf_demo/part1.wav" },
      { "url": "https://media.sume.com/artifacts/artf_demo/part2.wav" }
    ]
  }'

Will the joined track sound like one song?

Not automatically. Each generation is independent, so key, tempo and instruments can differ at the seam. Write the same tempo, key and instrument list into every prompt, give each part a distinct arc, and listen to the join. If you need one continuous composition of six minutes from a single model run, Sume's docs do not list a route for that.

What would a six-minute score cost on Sume?

Roughly three 2-minute tracks at $0.125 each, plus one Timeline audio join at $0.01, comes to $0.385, before any retries. The lengths are prompt targets, so you may need a fourth part or a trim. Timeline audio also has a split operation that cuts one file into up to 20 ranges, useful for trimming a part to fit.

That is a plan, not a benchmark: we did not generate a six-minute score for this post.

Should I generate the parts in order?

Generate them as independent takes, then order them. Because each part is a separate job, you can keep a strong first part and redo only a weak third one. Give every part a name in your notes, keep the accepted job ids, and join only after you have listened to each file. Timeline audio joins parts exactly as they are, with no crossfade, so the seams are only as smooth as the endings and openings you asked for.

Sources

Related posts

More in Models

All Models posts

Written by Sume