Suno Speech beta: voice and music in one take vs Sume jobs
Suno Speech (beta, Oct 1 2026) makes a spoken track and its soundtrack together. On Sume you chain a TTS job, a Music job and a Timeline render instead.

Suno's Speech beta, announced October 1, 2026, generates a spoken performance and its soundtrack as one creation. Sume does not have a single model that does both: you make the voice with a TTS job, the bed with a Music job, and combine them in a Timeline render with ducking.
That is more steps, but each piece is a separate artifact you can redo on its own.
What Suno says Speech does
Suno's release notes list Speech (Beta) under October 1, 2026 as a new model that creates speech and its soundtrack together, with examples such as bedtime stories over soft piano and hype speeches over stadium drums. The notes say it is on mobile and web and still in beta (read 2026-10-02).
| Date | Release | What the notes say |
|---|---|---|
| Oct 1, 2026 | Speech (beta) | Speech and soundtrack created together; mobile and web |
| Sep 17, 2026 | Suno Studio: improved MIDI | Faster MIDI to audio conversion; chat bar can create stems from a MIDI clip |
| Sep 9, 2026 | v6 family | v6, v6-wild and v6-mini; v6 and v6-wild for paid users |
The same result on Sume, in three jobs
Sume's docs describe separate surfaces for each piece, so the workflow is a chain rather than one prompt.
- Voice: a text-to-speech job (
POST /v1/tts-1.0/generate), up to 20,000 characters of transcript per request. - Music:
POST /v1/music-router/generatewith a 1 to 5000 character prompt. Say in the prompt that it is instrumental with no vocals, and describe the mood and instruments. - Mix: a Timeline 1.0 render where the voice file is the audio spine and the music file is the
soundtrack, withloop,gain_db,fade_out_secondsup to 10 andduck_dbfrom 0 to 20.
What you give up
Suno describes one model making both layers in a single take. On Sume the music model never hears the words, so timing is yours to manage: ask for a bed with a quiet middle section, then lower it under the voice with duck_db.
Timeline renders a video file. Its video slots take clips or stills, so for a voice-plus-music piece you supply at least a still. Sume does not return a single mixed audio file from this flow.
Speech beta pricing and plan limits were not on the pages read for this post, so compare cost yourself. Sume's Music price is a fixed amount per accepted generation ($0.125 on Music 1.0 per the docs), and TTS is billed by character.
A minimal Timeline body
Voice URL and music URL must both be Sume-hosted, which is why you import or generate first. A bed under a 30-second voice track looks like this.
{
"audio": { "url": "https://media.sume.com/artifacts/artf_demo/voice.wav", "duration_seconds": 30 },
"video": [ { "source_url": "https://media.sume.com/artifacts/artf_demo/still.png", "start": 0, "duration": 30 } ],
"soundtrack": { "url": "https://media.sume.com/artifacts/artf_demo/bed.mp3", "gain_db": -6, "loop": true, "fade_out_seconds": 3, "duck_db": 12 }
}Sources
Related posts
More in Comparisons
- Suno Studio MIDI update: AI music to MIDI vs Sume audio files
Suno Studio's Sept 17 2026 MIDI update converts MIDI to audio and makes stems from a MIDI clip. Sume's music API returns one audio file, with no MIDI in or out.
- Synthesia 3.0 clickable branching video: what Sume can do instead
Synthesia 3.0 lists clickable calls to action and branching. Sume renders linear avatar clips, so branches are separate clips that your own player links.
- Synthesia Assistant adds B-roll free: the itemized Sume way
Synthesia's Assistant adds curated AI B-roll at no extra cost (Sep 15, 2026). In Sume the avatar clip and B-roll are separate priced jobs you compose.
- Synthesia B-roll in 720p, 1080p or 4K vs Sume Avatar 1.0 at 720p
Synthesia lets you pick B-roll resolution (720p, 1080p, 4K) as of Sep 25, 2026. Sume avatar video renders at 720p today; here is how to add b-roll on top.
Written by Sume