Suno Speech makes voice and music as one track. Sume keeps them apart
Suno Speech beta renders voice and music together. On Sume you generate speech, a music bed and a ducked mix as separate jobs, so each can be redone alone.

Suno Speech is a beta model that generates a spoken voice and its background music as one finished track. Sume does not list a model that does this. Instead, you make the voice with TTS, the bed with the Music Router, and mix them in Timeline 1.0, so you can redo any one part without touching the others.
Suno announced Speech on October 1, 2026 and says it is opening the beta to everyone (Suno blog, read 2026-10-10). The decision is not which is better. It is whether you need to edit the voice and the music separately after the first render.
What Suno says Speech does
Suno describes Speech as the first audio model that generates voice and music together as one cohesive track. You type an idea, a poem or something you wrote, then describe the voice and the musical style. The output is spoken audio set to original background music.
The post does not state length limits, pricing, credits or commercial rights, so this article does not either. It also admits the beta is rough: Suno writes that British accents can wander off to Australia and back, and that dramatic pauses may be very dramatic.
| Question | Suno Speech beta | Sume |
|---|---|---|
| Output | One track: voice plus original background music | Separate files: a TTS voice file, a music file, and a rendered mix |
| How you describe it | Text plus a description of voice and musical style | Transcript to TTS; a separate music prompt (up to 5000 characters) to the Music Router |
| Change only the music | Not stated in the post | Regenerate the bed; the voice file stays as it is |
| Change only a spoken line | Not stated in the post | Regenerate that line and join with timeline audio concat |
| Music level under voice | Not stated in the post | Timeline soundtrack field duck_db, 0 to 20, plus gain_db |
| Price and length | Not stated in the post | Music: fixed $0.125 per generation; Timeline: $0.10 per output minute |
The Sume path, step by step
Each step is its own job with its own result, so a bad take costs you one step and not the whole track.
Generate the bed first. The Music Router takes a text prompt and picks the engine with sume/music-auto. Sume rejects duration fields, so write the length into the prompt, and put exclusions in the positive prompt such as Instrumental, no vocals, because a non-empty negative_prompt is refused.
- Voice: POST /v1/tts-1.0/generate with a transcript and an avatar_id or voice.id. Set language for any non-English text. Ask for wav when the file feeds a join.
- Music: POST /v1/music-router/generate with a prompt. Read the audio artifact from result.artifacts[] where type is audio.
- Import both files so they live on media.sume.com. Timeline accepts only this workspace's hosted media.
- Mix: POST /v1/timeline-1.0/render with the voice as the audio spine and the music as soundtrack with duck_db.
The mix request
The audio spine is the voice. The soundtrack field is the bed, and duck_db lowers it while the spine speaks. Duck needs a real spine, not silence, so a ducked bed always sits under actual speech. This request follows the Timeline 1.0 docs; replace the demo URLs with your own imported artifacts.
curl -X POST https://api.sume.com/v1/timeline-1.0/render \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: speech-mix-001" \
-d '{
"audio": {
"url": "https://media.sume.com/artifacts/artf_demo/voice.wav",
"duration_seconds": 24
},
"soundtrack": {
"url": "https://media.sume.com/artifacts/artf_demo/bed.mp3",
"gain_db": -6,
"duck_db": 12,
"fade_out_seconds": 3
},
"video": [
{
"source_url": "https://media.sume.com/artifacts/artf_demo/scene.mp4",
"start": 0,
"duration": 24
}
]
}'When one track is enough, and when to split
A single generated track suits a poem, a greeting or a quick spoken piece where the words and the music are one idea and you are happy to re-roll the whole thing. Sume has no equivalent single-prompt model, so that use case is outside what Sume ships today.
Split stems suit anything with a script that must stay exact, a brand voice that must stay the same across versions, or a bed you want to swap per channel. Long scripts can be sliced into takes and rejoined gaplessly with timeline audio concat (1 to 20 parts, $0.01 per job), and the returned segments[] offsets tell you where each line starts.
- Need the same voice with a different bed: regenerate music only.
- Need one corrected sentence: regenerate that take and concat again.
- Need the mix on a picture: the render already takes the clip and the audio together.
Sources
Related posts
More in Media tools
- TikTok ad 516 kbps bitrate floor: smallest file for 10 minutes
TikTok lists a minimum bitrate of 516 kbps and a 500 MB cap for non-Spark ads. The math gives 38.7 MB at the floor for 10 minutes, and Sume cannot tune bitrate.
- TikTok ad account name: 20 characters, and Spark caption: 4 lines
TikTok's In-Feed page lists a 20-character account name (10 for CJK scripts) and a 4-line Spark caption. Write them before you render the Sume clip.
- TikTok Ad Network thumbnail is the first frame: check frame 0
TikTok's Ad Network spec says the first frame is the default thumbnail. Pull frame 0 with Sume video frames, and skip a fade-in so the thumbnail is not black.
- TikTok ad AI disclosure: the AIGC label or your own caption
TikTok's ad policy accepts the AIGC label or your own clear disclaimer, caption, watermark or sticker. Here is what Sume can burn in and what it cannot.
Written by Sume