AI radio commercial generator: script, voice, music bed

An AI radio commercial is a timed script, an AI voice, and a music bed mixed into one audio file. How to generate each part with Sume's API.

5 min readSume
All posts

An AI radio commercial generator makes an audio ad in three parts: a script timed to the slot, an AI voice reading it, and a music bed mixed under the voice into one audio file. With Sume's API, text to speech voices the script, the Music Router makes an instrumental bed from a written brief, a Timeline 1.0 render mixes the two with the music ducked under the voice, and audio detach returns the finished spot as a WAV or MP3.

The steps follow Sume's Music Router, Music 1.0, Timeline 1.0, and Audio detach docs and the TTS schema in the Sume API reference, read on 2026-09-28; anything called current behavior is read from Sume's code, and prices from the code behind API pricing. Stations and ad platforms set their own rules for length, loudness, and file type, so check those with whoever will run the spot.

How many words fit in a 30-second radio spot?

It depends on the voice and the pace, so treat any count as a starting point. Sume's avatar-video planner estimates 2.8 words per second in current code; applied to a spot, that gives the rough budgets below. It is an estimate made for avatar clips, not a measured rate for any voice, and not a rule.

Then measure the real read. Ask text to speech for word timestamps, and in current code the result reports the file's duration_seconds. If the spot runs long or short, cut or add words, or set generation_config.speed, a multiplier from 0.6 to 1.5, and measure again; text to speech time calculator covers the fitting.

Word budgets at 2.8 words per second, the estimate in Sume's current avatar-video code; measure with the TTS fields in the Sume API reference, read 2026-09-28.
Spot lengthScript budget
15 secondsAbout 42 words
30 secondsAbout 84 words
60 secondsAbout 168 words

How do I voice the script?

Send it to POST /v1/tts-1.0/generate with a voice: the avatar_id or avatar_handle of an avatar whose voice is ready, or a voice.id such as a voice from your Voices library. Set language for a non-English script; omitted, it defaults to English. The mix inherits the voice file's sample rate and channel count, so voice it at the quality you want to keep. Use a cloned voice only with its owner's permission; AI voiceover in your own voice covers cloning. How to make an AI voiceover shows the full request.

How do I make the music bed?

Describe the bed in a prompt to POST /v1/music-router/generate. The request has no duration field, so put the length in the prompt. The music docs suggest a brief that names the emotion, the genre, a tempo in BPM, the key, two to four instruments, one named moment, and the era or production, then closes with “Instrumental, no vocals.”, adding “no spoken word” when a voice will sit on top. They call these creative directions, not guaranteed settings, so listen before you mix; if the bed comes out short, the mix can loop it.

curl -X POST https://api.sume.com/v1/music-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: radio-bed-001" \
  -d '{
    "prompt": "Bright, friendly acoustic pop, 104 BPM, G major. Strummed acoustic guitar, light hand claps, warm upright bass. Steady and uncluttered, with a small lift at 0:20. Clean modern production. A 30-second track. Instrumental, no vocals, no spoken word."
  }'

How do I mix the voice and the music into one spot?

Render once in Timeline 1.0 with the voice as audio.url, audio.duration_seconds set to the voice's length, the music as soundtrack with a duck_db dip under speech and a fade_out_seconds fade of up to 10 seconds, and one Sume-hosted still in video[]. Then send the MP4 to audio detach for a WAV (the default) or a 128 kbps MP3. The spot is as long as the voice file: the output runs audio.duration_seconds, and the bed fades over the last seconds of the spine. How to mix voice with background music has both calls.

What does an AI radio ad cost?

Four jobs, each billed plus a 5.5% agent fee by default:

  • Voice: $0.0475 per 1,000 characters, spaces and punctuation included. Each new take bills again.
  • Music bed: $0.125 per audio, the fixed price per generation.
  • Mix: $0.10 per output minute, reserved as ceil(audio.duration_seconds / 60) minutes, so a spot of 60 seconds or less reserves one minute.
  • Audio file: $0.01 per job, the rate on the audio detach docs page, which says to confirm it in GET /v1/catalog.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume