Suno Speech hype speech over drums: build it as two files on Sume
Suno Speech makes a speech and its soundtrack in one take. On Sume you make the voice and the drums as two files, then duck the music in Timeline.

Suno's Speech model makes a speech and its soundtrack in one take. Sume does not ship a single speech-plus-music model, so you build the same result from two files: a Sume TTS take for the voice, a Music Router take for the drums, and a Timeline soundtrack with duck_db to sit the drums under the words. The cost is two jobs and one render.
What Suno announced
Suno's release notes list Speech as a beta from October 1, 2026, described as the first model to create a speech and its soundtrack together, in one take. The examples on that page are a bedtime story over soft piano, a hype speech over stadium drums, and an ASMR grocery list. It is available on mobile and web.
The trade-off of one take is that the voice and the bed arrive mixed. If you want to change one word of the speech, or swap the drums for strings, you regenerate both. Two files let you change either one alone.
Step 1: the voice
Write the speech, pick a voice id, and send it to Sume TTS 1.0. The generation_config.emotion string (1 to 64 characters) and speed (0.6 to 1.5) carry the hype. Ask for a wav so the file stays sample-exact if you join it later.
TTS 1.0 bills $0.0475 per 1,000 characters on Sume, so a 1,000-character speech is $0.0475. The request maximum is 20,000 characters.
Step 2: the drums
Music Router bills a fixed $0.125 per generation. Write the brief with the seven axes: emotion, genre, BPM, key, two to four textured instruments, an arc with a named moment, and an era. End with the instrumental clause. Because the music sits under narration, also add **no spoken word** so the model does not add chants over your voice.
curl -X POST https://api.sume.com/v1/music-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: hype-drums-001" \
-d '{
"model": "sume/music-auto",
"prompt": "Triumphant arena rock, 112 BPM, E minor. Stadium kick and snare, low toms, a rising snare roll at 0:20 into a crash. A 1-minute track. Instrumental, no vocals, no spoken word."
}'Step 3: duck and render
In a Timeline 1.0 render, send the speech as the audio spine and the drums as soundtrack. The soundtrack object takes url, gain_db, loop, fade_out_seconds (up to 10) and duck_db (0 to 20). Ducking needs a real spine, so the voice file must be the audio.url. See the duck walkthrough for the render body.
Set duck_db to a value in range and listen once. The drums should dip under each sentence and return between them, which is the effect the one-take model aims for.
When Suno's one take is the better fit
If you want a finished clip where the model decides how the music reacts to the words, and you do not need to edit either part afterwards, one take is shorter. Choose the two-file route when the script changes often, when you need word timestamps for captions, or when the same voice must read ten scripts over different beds.
- Edit one word: re-run only the TTS job.
- Swap the bed: re-run only the music job, $0.125.
- Captions: ask TTS for
timestamps: {words: true}and keep the words for captions.
Budget the two-file route
A one-minute hype speech is roughly 800 to 1,000 characters, so the voice is about four to five cents and the drums are $0.125. Add the Timeline render on top. If you want three beds to compare, three music takes are $0.375, and the voice file is reused for all of them.
That is the practical upside of the split: you audition the cheap part (music) many times against one fixed voice take.
Sources
Related posts
More in Use cases
- Suno v6 on licensed data: does it change brand commercial-use risk
Suno v6 is reported to use licensed training data, but suits continue and commercial rights still need a paid plan. What changes for a brand and what does not.
- Suno retires older models: keeping a brand jingle reproducible
TechCrunch says Suno will retire older models after v6. A brand jingle needs a stored master file and a saved prompt, not a promise to regenerate it.
- Swap the person in a clip: reference video or H3 Max Recast on Sume
Seedance 2.5 offers reference-based editing. For a straight person swap, Sume lists H3 Max Recast. When to use which, and the inputs each needs.
- Replace audio from 12.4 s to 14 s of a clip with Timeline parts
Seedance 2.5 lists timestamp-level audio editing. On Sume you can detach a clip's audio and rebuild it from parts with a replaced line in the middle.
Written by Sume