Suno Speech hype speech over drums: build it as two files on Sume

Suno Speech makes a speech and its soundtrack in one take. On Sume you make the voice and the drums as two files, then duck the music in Timeline.

5 min readSume
All posts

Suno's Speech model makes a speech and its soundtrack in one take. Sume does not ship a single speech-plus-music model, so you build the same result from two files: a Sume TTS take for the voice, a Music Router take for the drums, and a Timeline soundtrack with duck_db to sit the drums under the words. The cost is two jobs and one render.

What Suno announced

Suno's release notes list Speech as a beta from October 1, 2026, described as the first model to create a speech and its soundtrack together, in one take. The examples on that page are a bedtime story over soft piano, a hype speech over stadium drums, and an ASMR grocery list. It is available on mobile and web.

The trade-off of one take is that the voice and the bed arrive mixed. If you want to change one word of the speech, or swap the drums for strings, you regenerate both. Two files let you change either one alone.

Step 1: the voice

Write the speech, pick a voice id, and send it to Sume TTS 1.0. The generation_config.emotion string (1 to 64 characters) and speed (0.6 to 1.5) carry the hype. Ask for a wav so the file stays sample-exact if you join it later.

TTS 1.0 bills $0.0475 per 1,000 characters on Sume, so a 1,000-character speech is $0.0475. The request maximum is 20,000 characters.

Step 2: the drums

Music Router bills a fixed $0.125 per generation. Write the brief with the seven axes: emotion, genre, BPM, key, two to four textured instruments, an arc with a named moment, and an era. End with the instrumental clause. Because the music sits under narration, also add **no spoken word** so the model does not add chants over your voice.

curl -X POST https://api.sume.com/v1/music-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: hype-drums-001" \
  -d '{
    "model": "sume/music-auto",
    "prompt": "Triumphant arena rock, 112 BPM, E minor. Stadium kick and snare, low toms, a rising snare roll at 0:20 into a crash. A 1-minute track. Instrumental, no vocals, no spoken word."
  }'

Step 3: duck and render

In a Timeline 1.0 render, send the speech as the audio spine and the drums as soundtrack. The soundtrack object takes url, gain_db, loop, fade_out_seconds (up to 10) and duck_db (0 to 20). Ducking needs a real spine, so the voice file must be the audio.url. See the duck walkthrough for the render body.

Set duck_db to a value in range and listen once. The drums should dip under each sentence and return between them, which is the effect the one-take model aims for.

When Suno's one take is the better fit

If you want a finished clip where the model decides how the music reacts to the words, and you do not need to edit either part afterwards, one take is shorter. Choose the two-file route when the script changes often, when you need word timestamps for captions, or when the same voice must read ten scripts over different beds.

  • Edit one word: re-run only the TTS job.
  • Swap the bed: re-run only the music job, $0.125.
  • Captions: ask TTS for timestamps: {words: true} and keep the words for captions.

Budget the two-file route

A one-minute hype speech is roughly 800 to 1,000 characters, so the voice is about four to five cents and the drums are $0.125. Add the Timeline render on top. If you want three beds to compare, three music takes are $0.375, and the voice file is reused for all of them.

That is the practical upside of the split: you audition the cheap part (music) many times against one fixed voice take.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume