How to make a pep talk audio track with music: voice, bed and render

A 90-second pep talk with a music bed is one TTS job, one Music Router job and one render: about $0.39 on Sume. Suno Speech beta does it in one pass.

4 min readSume
All posts

A 90-second pep talk with music is three Sume jobs: a text-to-speech job for the voice (about 1,125 characters), a Music Router job for the bed, and a Timeline render that lays the bed under the voice. The bill is $0.06 for the voice, $0.125 for the track and $0.20 for the render, $0.385 in total at the rates in the catalog on 2026-10-07. Suno's Speech beta, opened to everyone on 2026-10-01, produces voice and original music as one track from one prompt; on Sume the same result is assembled from separate files, which costs a few extra calls and gives you control over each layer.

Suno's announcement lists pep talks next to meditations, stories and dramatic readings as the kind of thing Speech is for, and says nothing about price or languages. That is a fair reason to build the same piece from parts you can price today.

Write the script for the clock

Cartesia's pricing page puts one minute of Sonic speech at 750 to 800 credits, one credit per character, and Sume TTS 1.0 meters one credit per character on the same engine. At 750 characters a minute, 90 seconds of speech is 1,125 characters, spaces and punctuation included. A pep talk reads best with short sentences and some white space, so aim a little under that and let the pauses carry the length.

At $0.0475 per 1,000 characters, 1,125 characters is $0.0534, which the catalog rounds up to $0.06. Redoing a take costs the same again, so listen to the first 200 characters on a cheap trial job before committing the full script.

Voice job: pace and emotion

TTS 1.0 takes an optional generation_config with speed from 0.6 to 1.5, volume from 0.5 to 2.0 and a free-text emotion guide. For a pep talk, leave speed at 1.0 or nudge it to 1.05 and describe the delivery in the emotion string. Request wav so the file can be reused in a render without re-encoding padding.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: pep-talk-voice-001" \
  -d @- <<JSON
{
  "transcript": "You showed up. That is the hard part. Now take one more step.",
  "voice": { "id": "$VOICE_ID" },
  "language": "en",
  "generation_config": { "speed": 1.05, "emotion": "urgent, warm, building" },
  "output_format": { "container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100 },
  "mode": "async"
}
JSON

Music job: leave room for the voice

Music Router takes a prompt of 1 to 5,000 characters, rejects duration, and expects the length inside the prompt. Say how long the track is, name a tempo, and end with 'Instrumental, no vocals' so the bed does not compete with the speaker. Put exclusions in the positive prompt; a non-empty negative_prompt returns a 400.

A prompt that works for this use: 'A 90-second driving cinematic build, 118 BPM, E minor. Pulsing synth bass, tight kick, rising strings, a drop back to bass and claps at 0:45, full return at 1:05. Instrumental, no vocals.' Every Music Router model charges the same fixed price per generation, $0.125.

Render: bed under the voice

Timeline 1.0 takes the voice file as the audio spine and the track as soundtrack, with gain_db, loop, fade_out_seconds up to 10 and duck_db from 0 to 20 to pull the bed down under speech. The render needs at least one video slot, so a still image held for the full length works; the output is an MP4 at $0.10 per started output minute, and 90 seconds starts a second minute, so $0.20. Sume documents no audio-only mixdown, so if you need an audio file, keep the voice and the bed as two files and mix in your own editor.

curl -X POST https://api.sume.com/v1/timeline-1.0/render \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: pep-talk-render-001" \
  -d '{
    "audio": { "url": "https://media.sume.com/artifacts/artf_demo/voice.wav", "duration_seconds": 90 },
    "video": [{ "source_url": "https://media.sume.com/artifacts/artf_demo/cover.png", "start": 0, "duration": 90 }],
    "soundtrack": { "url": "https://media.sume.com/artifacts/artf_demo/bed.mp3", "gain_db": -6, "duck_db": 10, "fade_out_seconds": 4 }
  }'

What it costs, line by line

Where Suno Speech gives you one finished track and a single regenerate button, the split approach lets you swap the bed without touching the voice. That matters for a pep talk, where the words are usually approved first and the music is the part you keep tuning. Keep the voice job id and the music job id in your notes so either layer can be fetched again.

  • If you drop the render and mix elsewhere, the audio costs $0.185 for both files.
  • Retake only the layer that failed: a new voice take is $0.06, a new bed is $0.125.
Cost of a 90-second pep talk with a music bed on Sume, rates as of 2026-10-07 (catalog entries for TTS 1.0, Music Router and Timeline 1.0); vendor figures read 2026-10-07.
StepUnitPriceThis piece
TTS 1.0 voiceper 1,000 characters$0.0475$0.06
Music Router trackper generation$0.125$0.125
Timeline 1.0 renderper started output minute$0.10$0.20
Total$0.385

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume