How to make a pep talk audio track with music: voice, bed and render
A 90-second pep talk with a music bed is one TTS job, one Music Router job and one render: about $0.39 on Sume. Suno Speech beta does it in one pass.

A 90-second pep talk with music is three Sume jobs: a text-to-speech job for the voice (about 1,125 characters), a Music Router job for the bed, and a Timeline render that lays the bed under the voice. The bill is $0.06 for the voice, $0.125 for the track and $0.20 for the render, $0.385 in total at the rates in the catalog on 2026-10-07. Suno's Speech beta, opened to everyone on 2026-10-01, produces voice and original music as one track from one prompt; on Sume the same result is assembled from separate files, which costs a few extra calls and gives you control over each layer.
Suno's announcement lists pep talks next to meditations, stories and dramatic readings as the kind of thing Speech is for, and says nothing about price or languages. That is a fair reason to build the same piece from parts you can price today.
Write the script for the clock
Cartesia's pricing page puts one minute of Sonic speech at 750 to 800 credits, one credit per character, and Sume TTS 1.0 meters one credit per character on the same engine. At 750 characters a minute, 90 seconds of speech is 1,125 characters, spaces and punctuation included. A pep talk reads best with short sentences and some white space, so aim a little under that and let the pauses carry the length.
At $0.0475 per 1,000 characters, 1,125 characters is $0.0534, which the catalog rounds up to $0.06. Redoing a take costs the same again, so listen to the first 200 characters on a cheap trial job before committing the full script.
Voice job: pace and emotion
TTS 1.0 takes an optional generation_config with speed from 0.6 to 1.5, volume from 0.5 to 2.0 and a free-text emotion guide. For a pep talk, leave speed at 1.0 or nudge it to 1.05 and describe the delivery in the emotion string. Request wav so the file can be reused in a render without re-encoding padding.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: pep-talk-voice-001" \
-d @- <<JSON
{
"transcript": "You showed up. That is the hard part. Now take one more step.",
"voice": { "id": "$VOICE_ID" },
"language": "en",
"generation_config": { "speed": 1.05, "emotion": "urgent, warm, building" },
"output_format": { "container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100 },
"mode": "async"
}
JSONMusic job: leave room for the voice
Music Router takes a prompt of 1 to 5,000 characters, rejects duration, and expects the length inside the prompt. Say how long the track is, name a tempo, and end with 'Instrumental, no vocals' so the bed does not compete with the speaker. Put exclusions in the positive prompt; a non-empty negative_prompt returns a 400.
A prompt that works for this use: 'A 90-second driving cinematic build, 118 BPM, E minor. Pulsing synth bass, tight kick, rising strings, a drop back to bass and claps at 0:45, full return at 1:05. Instrumental, no vocals.' Every Music Router model charges the same fixed price per generation, $0.125.
Render: bed under the voice
Timeline 1.0 takes the voice file as the audio spine and the track as soundtrack, with gain_db, loop, fade_out_seconds up to 10 and duck_db from 0 to 20 to pull the bed down under speech. The render needs at least one video slot, so a still image held for the full length works; the output is an MP4 at $0.10 per started output minute, and 90 seconds starts a second minute, so $0.20. Sume documents no audio-only mixdown, so if you need an audio file, keep the voice and the bed as two files and mix in your own editor.
curl -X POST https://api.sume.com/v1/timeline-1.0/render \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: pep-talk-render-001" \
-d '{
"audio": { "url": "https://media.sume.com/artifacts/artf_demo/voice.wav", "duration_seconds": 90 },
"video": [{ "source_url": "https://media.sume.com/artifacts/artf_demo/cover.png", "start": 0, "duration": 90 }],
"soundtrack": { "url": "https://media.sume.com/artifacts/artf_demo/bed.mp3", "gain_db": -6, "duck_db": 10, "fade_out_seconds": 4 }
}'What it costs, line by line
Where Suno Speech gives you one finished track and a single regenerate button, the split approach lets you swap the bed without touching the voice. That matters for a pep talk, where the words are usually approved first and the music is the part you keep tuning. Keep the voice job id and the music job id in your notes so either layer can be fetched again.
- If you drop the render and mix elsewhere, the audio costs $0.185 for both files.
- Retake only the layer that failed: a new voice take is $0.06, a new bed is $0.125.
| Step | Unit | Price | This piece |
|---|---|---|---|
| TTS 1.0 voice | per 1,000 characters | $0.0475 | $0.06 |
| Music Router track | per generation | $0.125 | $0.125 |
| Timeline 1.0 render | per started output minute | $0.10 | $0.20 |
| Total | $0.385 |
Sources
Related posts
More in Use cases
- Pick clip lengths from a voice-over script: one sentence per clip
Turn a voice-over into clips: measure each voiced sentence, round up, and pick the models whose duration window contains it. Windows for six Sume models.
- How do I turn one podcast episode into five quote clips for social?
Transcribe the episode in 10-minute chunks, split five quotes out, put the cover still under each and burn captions: $1.80 for a 28-minute episode on Sume.
- Post-call recap video with an AI avatar for prospects: script and cost
After a sales call, send a 30-second recap clip from a Sume avatar. Script structure, cost at standard, plus and max, and how to keep it honest and reviewed.
- Pre-rendered avatar greetings per visitor segment, not a live avatar
Instead of a live avatar for each visitor, render one short Sume avatar clip per segment ahead of time. Per-tier cost for six 12-second greetings.
Written by Sume