Suno Speech makes voice and music in one pass. What does Sume do?

Suno's Speech beta (Oct 1, 2026) makes voice and music as one track. Sume has no such model; it layers TTS, Music and a Timeline soundtrack. Trade-offs inside.

4 min readSume
All posts

Sume does not have a model that generates spoken voice and music together as one track, so it cannot match Suno's new Speech beta in a single call. What Sume offers is a three-step stack: Sume TTS for the voice, Music 1.0 for an instrumental bed, and a Timeline 1.0 soundtrack that mixes them with ducking.

Suno announced Speech (beta) on October 1, 2026, and the post was read on 2026-10-10. This page compares the two approaches without ranking audio quality, which neither vendor page lets us judge.

What Suno says Speech is

Suno's blog describes Speech as the first audio model that generates voice and music together as one cohesive track. It is a beta built into the Suno platform. The post does not mention an API, a commercial-use term or a length limit, so none of those can be assumed.

It also lists known issues in its own words: British accents can wander off to Australia and back, and dramatic pauses may be very dramatic. Those are useful signals for anyone planning to ship a read-through without listening to it first.

The Sume stack, step by step

The voice comes from TTS 1.0 at $0.0475 per 1,000 characters, up to 20,000 characters per request. You can set speed between 0.6 and 1.5, volume between 0.5 and 2, and an emotion in generation_config. Timestamps for words are available if you need them.

The bed comes from Music 1.0 at a flat $0.125 per generation. There is no duration parameter; a duration or a non-empty negative_prompt is rejected, so length is steered in the prompt and trimmed later. To keep a bed from singing over your narrator, say so in the prompt, for example by asking for an instrumental with no vocals and no spoken word. The music router lists Lyria 3.5 and Lyria 3 Pro as selectable ids.

The mix happens in Timeline 1.0. Its soundtrack object takes a url, gain_db, loop, a fade_out_seconds of up to 10 and a duck_db of 0 to 20 that lowers the bed under speech; ducking needs a real audio spine, so the voice file is the spine. Timeline render is priced at $0.10 per output minute.

Side by side

The table separates what the Suno post says from what the Sume docs and code show.

One-pass voice and music vs a layered Sume stack; Suno column from its blog read 2026-10-10, Sume column from docs and code
QuestionSuno Speech (beta)Sume stack
Calls to makeOne generationThree jobs: TTS, Music, Timeline
Control of the voice aloneNot describedYes, TTS settings and a separate file
Swap the music onlyNot describedYes, regenerate the bed and re-render
APINone mentioned in the postREST, async jobs, webhooks
Price shownNot in the post$0.0475 per 1,000 chars, $0.125 per track, $0.10 per output minute
Known caveatsAccent drift, long pausesBed length is prompt-steered; no loudness normalizer

Which one fits which job

A one-pass model suits a quick demo or a jingle where the voice and the music are one creative object. A layered stack suits work you will revise: if the client changes one sentence, you regenerate only that line of TTS, re-join, and re-render, instead of regenerating a whole song.

Costs favor the stack for short reads. A 1,000-character script is $0.0475 of voice, $0.125 of music and, for a 60-second mix, $0.10 of timeline, so $0.2725 in total at list prices. Add video slots and the timeline minute rate still applies per output minute.

For the mixing details, see ducking music under voiceover and trimming music to an exact length.

What to verify before you choose

Read Suno's current terms for commercial use before you build on Speech, since the announcement post is silent on it. On the Sume side, confirm the rate in the live catalog, and test your voice and bed together: the stack gives control, but it does not decide taste for you.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume