AI radio commercial generator: script, voice, music bed
An AI radio commercial is a timed script, an AI voice, and a music bed mixed into one audio file. How to generate each part with Sume's API.

An AI radio commercial generator makes an audio ad in three parts: a script timed to the slot, an AI voice reading it, and a music bed mixed under the voice into one audio file. With Sume's API, text to speech voices the script, the Music Router makes an instrumental bed from a written brief, a Timeline 1.0 render mixes the two with the music ducked under the voice, and audio detach returns the finished spot as a WAV or MP3.
The steps follow Sume's Music Router, Music 1.0, Timeline 1.0, and Audio detach docs and the TTS schema in the Sume API reference, read on 2026-09-28; anything called current behavior is read from Sume's code, and prices from the code behind API pricing. Stations and ad platforms set their own rules for length, loudness, and file type, so check those with whoever will run the spot.
How many words fit in a 30-second radio spot?
It depends on the voice and the pace, so treat any count as a starting point. Sume's avatar-video planner estimates 2.8 words per second in current code; applied to a spot, that gives the rough budgets below. It is an estimate made for avatar clips, not a measured rate for any voice, and not a rule.
Then measure the real read. Ask text to speech for word timestamps, and in current code the result reports the file's duration_seconds. If the spot runs long or short, cut or add words, or set generation_config.speed, a multiplier from 0.6 to 1.5, and measure again; text to speech time calculator covers the fitting.
| Spot length | Script budget |
|---|---|
| 15 seconds | About 42 words |
| 30 seconds | About 84 words |
| 60 seconds | About 168 words |
How do I voice the script?
Send it to POST /v1/tts-1.0/generate with a voice: the avatar_id or avatar_handle of an avatar whose voice is ready, or a voice.id such as a voice from your Voices library. Set language for a non-English script; omitted, it defaults to English. The mix inherits the voice file's sample rate and channel count, so voice it at the quality you want to keep. Use a cloned voice only with its owner's permission; AI voiceover in your own voice covers cloning. How to make an AI voiceover shows the full request.
How do I make the music bed?
Describe the bed in a prompt to POST /v1/music-router/generate. The request has no duration field, so put the length in the prompt. The music docs suggest a brief that names the emotion, the genre, a tempo in BPM, the key, two to four instruments, one named moment, and the era or production, then closes with “Instrumental, no vocals.”, adding “no spoken word” when a voice will sit on top. They call these creative directions, not guaranteed settings, so listen before you mix; if the bed comes out short, the mix can loop it.
curl -X POST https://api.sume.com/v1/music-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: radio-bed-001" \
-d '{
"prompt": "Bright, friendly acoustic pop, 104 BPM, G major. Strummed acoustic guitar, light hand claps, warm upright bass. Steady and uncluttered, with a small lift at 0:20. Clean modern production. A 30-second track. Instrumental, no vocals, no spoken word."
}'How do I mix the voice and the music into one spot?
Render once in Timeline 1.0 with the voice as audio.url, audio.duration_seconds set to the voice's length, the music as soundtrack with a duck_db dip under speech and a fade_out_seconds fade of up to 10 seconds, and one Sume-hosted still in video[]. Then send the MP4 to audio detach for a WAV (the default) or a 128 kbps MP3. The spot is as long as the voice file: the output runs audio.duration_seconds, and the bed fades over the last seconds of the spine. How to mix voice with background music has both calls.
What does an AI radio ad cost?
Four jobs, each billed plus a 5.5% agent fee by default:
- Voice: $0.0475 per 1,000 characters, spaces and punctuation included. Each new take bills again.
- Music bed: $0.125 per audio, the fixed price per generation.
- Mix: $0.10 per output minute, reserved as
ceil(audio.duration_seconds / 60)minutes, so a spot of 60 seconds or less reserves one minute. - Audio file: $0.01 per job, the rate on the audio detach docs page, which says to confirm it in
GET /v1/catalog.
Sources
Related posts
More in Use cases
- AI Spotify Canvas generator: make a 3–8 second loop
Spotify asks for a 3–8 second, 9:16 Canvas, 720–1080 px tall. Generate a looping 9:16 clip with AI, then render it at an exact size and length.
- AI sprite generator: make 2D game sprites and sprite sheets
Make 2D game sprites with AI: one pose per image on a plain background, the same character reference every time, then a transparent PNG cutout.
- AI sticker generator: from a prompt or photo to a PNG cutout
Make stickers with AI: generate one flat design with a bold outline on a plain background, size it for print, then cut it out to a transparent PNG.
- AI T-shirt design generator: art on a transparent PNG
Generate T-shirt art on a plain background, upscale it to your printer's pixel size, then cut it out so the design sits on a transparent PNG.
Written by Sume