Poem reading video: slow TTS, an emotion string and a ducked music bed
Set generation_config.speed to 0.6 and an emotion string on the TTS router, add a Music Router bed, and render with soundtrack duck_db so the voice stays clear.

Set generation_config.speed to 0.6, the slowest the TTS router allows, add an emotion string, and generate the music bed separately. Then render with a soundtrack that has duck_db set, so the bed drops under the voice. The three jobs cost a few cents of speech plus $0.125 for music.
Step 1: the slow reading
POST /v1/tts-router/generate requires model. The generation_config object takes speed from 0.6 to 1.5, volume from 0.5 to 2 and an emotion string of up to 64 characters. The voice comes from voice.id, which must be a UUID or a voi_ id, or from avatar_id or avatar_handle.
curl -X POST https://api.sume.com/v1/tts-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: poem-voice-001" \
-d '{
"model": "sonic-3.6",
"transcript": "I wandered lonely as a cloud, that floats on high o'\''er vales and hills.",
"voice": { "id": "'"$SUME_VOICE_ID"'" },
"generation_config": { "speed": 0.6, "emotion": "calm" }
}'Step 2: the bed and the mix
Ask the Music Router for a quiet, sparse instrumental, for example Solo piano, slow, sparse, instrumental, no vocals. A 1-minute track. Length is steered in the prompt. In the Timeline 1.0 render, pass the speech as audio.url and the music as soundtrack.
| Setting | Range or value | Where |
|---|---|---|
| speed | 0.6 to 1.5 | TTS generation_config |
| emotion | String, up to 64 characters | TTS generation_config |
| duck_db | 0 to 20, needs a real audio spine | Timeline soundtrack |
| fade_out_seconds | Up to 10 | Timeline soundtrack |
| Music price | $0.125 per generation | Music Router |
Checks before you ship
Listen at 0.6 speed, because very slow delivery can sound stilted on short lines, and raise it a notch if it does. Keep duck_db modest. A poem needs the bed audible in the gaps. If the song is shorter than the spoken track, set loop on the soundtrack and a fade_out_seconds of a few seconds. The bedtime story post uses the same structure.
Sources
Related posts
More in Use cases
- Pre-render voice agent greetings and fixed replies as TTS files
Live TTS latency is wasted on lines that never change. Render the greeting, hold message and confirmations once as files with Sume TTS and play them instantly.
- Can you premiere a YouTube Short or make it members-only?
No Premieres for Shorts, but members-only Shorts exist: 60 seconds, square or vertical, original audio only. What Help says and how to cut a 60-second version.
- Price plate on 10 product clips for Black Friday: $0.40 total
Cutting ten product clips to 8 seconds and overlaying a price plate still costs $0.04 per SKU on Sume, or $0.40 for ten, ahead of Black Friday on Nov 27, 2026.
- Printful print file requirements and what to ask Sume for
Printful's print-file page: PNG or JPEG, 300 DPI ideal, 75 DPI minimum, sRGB, full bleed. The Sume Image API settings that fit.
Written by Sume