Suno Speech beta: voice and music in one take vs Sume jobs

Suno Speech (beta, Oct 1 2026) makes a spoken track and its soundtrack together. On Sume you chain a TTS job, a Music job and a Timeline render instead.

5 min readSume
All posts

Suno's Speech beta, announced October 1, 2026, generates a spoken performance and its soundtrack as one creation. Sume does not have a single model that does both: you make the voice with a TTS job, the bed with a Music job, and combine them in a Timeline render with ducking.

That is more steps, but each piece is a separate artifact you can redo on its own.

What Suno says Speech does

Suno's release notes list Speech (Beta) under October 1, 2026 as a new model that creates speech and its soundtrack together, with examples such as bedtime stories over soft piano and hype speeches over stadium drums. The notes say it is on mobile and web and still in beta (read 2026-10-02).

Suno release notes, Sept to Oct 2026 (read 2026-10-02)
DateReleaseWhat the notes say
Oct 1, 2026Speech (beta)Speech and soundtrack created together; mobile and web
Sep 17, 2026Suno Studio: improved MIDIFaster MIDI to audio conversion; chat bar can create stems from a MIDI clip
Sep 9, 2026v6 familyv6, v6-wild and v6-mini; v6 and v6-wild for paid users

The same result on Sume, in three jobs

Sume's docs describe separate surfaces for each piece, so the workflow is a chain rather than one prompt.

  • Voice: a text-to-speech job (POST /v1/tts-1.0/generate), up to 20,000 characters of transcript per request.
  • Music: POST /v1/music-router/generate with a 1 to 5000 character prompt. Say in the prompt that it is instrumental with no vocals, and describe the mood and instruments.
  • Mix: a Timeline 1.0 render where the voice file is the audio spine and the music file is the soundtrack, with loop, gain_db, fade_out_seconds up to 10 and duck_db from 0 to 20.

What you give up

Suno describes one model making both layers in a single take. On Sume the music model never hears the words, so timing is yours to manage: ask for a bed with a quiet middle section, then lower it under the voice with duck_db.

Timeline renders a video file. Its video slots take clips or stills, so for a voice-plus-music piece you supply at least a still. Sume does not return a single mixed audio file from this flow.

Speech beta pricing and plan limits were not on the pages read for this post, so compare cost yourself. Sume's Music price is a fixed amount per accepted generation ($0.125 on Music 1.0 per the docs), and TTS is billed by character.

A minimal Timeline body

Voice URL and music URL must both be Sume-hosted, which is why you import or generate first. A bed under a 30-second voice track looks like this.

{
  "audio": { "url": "https://media.sume.com/artifacts/artf_demo/voice.wav", "duration_seconds": 30 },
  "video": [ { "source_url": "https://media.sume.com/artifacts/artf_demo/still.png", "start": 0, "duration": 30 } ],
  "soundtrack": { "url": "https://media.sume.com/artifacts/artf_demo/bed.mp3", "gain_db": -6, "loop": true, "fade_out_seconds": 3, "duck_db": 12 }
}

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume