Suno Speech makes voice and music in one pass. What does Sume do?
Suno's Speech beta (Oct 1, 2026) makes voice and music as one track. Sume has no such model; it layers TTS, Music and a Timeline soundtrack. Trade-offs inside.

Sume does not have a model that generates spoken voice and music together as one track, so it cannot match Suno's new Speech beta in a single call. What Sume offers is a three-step stack: Sume TTS for the voice, Music 1.0 for an instrumental bed, and a Timeline 1.0 soundtrack that mixes them with ducking.
Suno announced Speech (beta) on October 1, 2026, and the post was read on 2026-10-10. This page compares the two approaches without ranking audio quality, which neither vendor page lets us judge.
What Suno says Speech is
Suno's blog describes Speech as the first audio model that generates voice and music together as one cohesive track. It is a beta built into the Suno platform. The post does not mention an API, a commercial-use term or a length limit, so none of those can be assumed.
It also lists known issues in its own words: British accents can wander off to Australia and back, and dramatic pauses may be very dramatic. Those are useful signals for anyone planning to ship a read-through without listening to it first.
The Sume stack, step by step
The voice comes from TTS 1.0 at $0.0475 per 1,000 characters, up to 20,000 characters per request. You can set speed between 0.6 and 1.5, volume between 0.5 and 2, and an emotion in generation_config. Timestamps for words are available if you need them.
The bed comes from Music 1.0 at a flat $0.125 per generation. There is no duration parameter; a duration or a non-empty negative_prompt is rejected, so length is steered in the prompt and trimmed later. To keep a bed from singing over your narrator, say so in the prompt, for example by asking for an instrumental with no vocals and no spoken word. The music router lists Lyria 3.5 and Lyria 3 Pro as selectable ids.
The mix happens in Timeline 1.0. Its soundtrack object takes a url, gain_db, loop, a fade_out_seconds of up to 10 and a duck_db of 0 to 20 that lowers the bed under speech; ducking needs a real audio spine, so the voice file is the spine. Timeline render is priced at $0.10 per output minute.
Side by side
The table separates what the Suno post says from what the Sume docs and code show.
| Question | Suno Speech (beta) | Sume stack |
|---|---|---|
| Calls to make | One generation | Three jobs: TTS, Music, Timeline |
| Control of the voice alone | Not described | Yes, TTS settings and a separate file |
| Swap the music only | Not described | Yes, regenerate the bed and re-render |
| API | None mentioned in the post | REST, async jobs, webhooks |
| Price shown | Not in the post | $0.0475 per 1,000 chars, $0.125 per track, $0.10 per output minute |
| Known caveats | Accent drift, long pauses | Bed length is prompt-steered; no loudness normalizer |
Which one fits which job
A one-pass model suits a quick demo or a jingle where the voice and the music are one creative object. A layered stack suits work you will revise: if the client changes one sentence, you regenerate only that line of TTS, re-join, and re-render, instead of regenerating a whole song.
Costs favor the stack for short reads. A 1,000-character script is $0.0475 of voice, $0.125 of music and, for a 60-second mix, $0.10 of timeline, so $0.2725 in total at list prices. Add video slots and the timeline minute rate still applies per output minute.
For the mixing details, see ducking music under voiceover and trimming music to an exact length.
What to verify before you choose
Read Suno's current terms for commercial use before you build on Speech, since the announcement post is silent on it. On the Sume side, confirm the rate in the live catalog, and test your voice and bed together: the stack gives control, but it does not decide taste for you.
Sources
Related posts
More in Comparisons
- Suno v6-mini for all users vs Sume Music at $0.125 a track
Suno v6-mini is open to every user, while v6 and v6-wild need Pro or Premier. Sume does not list Suno; it sells one Lyria track for a flat $0.125.
- Synthesia API rate limits by tier vs Sume's queue and 429s
Synthesia caps writes at 60 to 120 a minute by tier and answers 429. Sume queues accepted jobs and returns queue_full or rate_limited. Read 2026-10-10.
- Synthesia-Signature vs Sume's webhook signature: one verifier?
Synthesia signs timestamp.request_body with HMAC-SHA256; Sume signs the same shape. Header names, prefix and tolerance differ. Compare both (read 2026-10-10).
- Talking-video API status values: HeyGen, Synthesia, D-ID, Sume mapped
Four avatar-video APIs, four status vocabularies: pending, in_progress, started, queued. Map them to one state machine in your code (read 2026-10-10).
Written by Sume