AI voiceover for ads: one brand voice across every cut
Make AI voiceover for ads by pinning one voice and voicing each 6, 15 or 30 second cut as its own text to speech request. Fields, fit and cost.

AI voiceover for ads works by picking one voice for the brand and generating each cut's script, whether a 6-, 15- or 30-second version, as its own text to speech request in that voice. You measure the real length of each read, then cut picture to it, and a new variant is a new request rather than a new recording session.
On Sume that request is POST /v1/tts-1.0/generate, billed at $0.0475 per 1,000 characters plus a 5.5% agent fee by default. The fields below come from the TTS schema in the Sume API reference (see API reference), the rate from API pricing, and the mixing step from Timeline 1.0, read on 2026-09-29.
How do I keep the same voice across every ad variant?
Send the same voice selector on every request. There are two: voice.id, which takes a TTS voice UUID or a Voices library id (voi_ plus 32 hex), or the top-level avatar_id / avatar_handle, which uses a ready avatar's voice. If you send both, they must match, or the request fails with 400. A brand voice you clone or describe is made in the Sume app under Assets → Voices, not through the API; see AI voiceover in your own voice.
Keep the other settings fixed across the campaign too, so only the script changes between cuts:
| Field | What it does for an ad |
|---|---|
voice.id or avatar_id | The brand voice; same value on every cut |
language | Set it for every non-English script; omitted means English |
generation_config.speed | Pace multiplier from 0.6 to 1.5 |
generation_config.volume | Volume multiplier from 0.5 to 2.0 |
generation_config.emotion | A free-form emotion guide, up to 64 characters |
output_format | Defaults to MP3 at 44,100 Hz and 128 kbps; WAV is available |
timestamps.words | Returns word timings on the result |
transcript | The script, up to 20,000 characters |
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: spring-sale-15s-v1" \
-d '{
"transcript": "Spring sale starts Friday. Everything in store, twenty percent off. Visit example.com.",
"voice": { "mode": "id", "id": "9f3e2b1c-6a4d-4c8e-9f1a-2b3c4d5e6f70" },
"generation_config": { "speed": 1.1 },
"output_format": { "container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100 },
"timestamps": { "words": true },
"mode": "async"
}'How do I make a read fit a 6, 15 or 30 second slot?
Write to the slot, then measure. With timestamps.words set, the completed result carries words[] with start and end seconds, so the last word's end is the spoken length. If a read runs long, cut words first and nudge generation_config.speed second. The estimate-then-measure method is in Text to speech time calculator.
How do I put the voiceover under the ad's picture?
In a Timeline 1.0 render, the voiceover file becomes the audio spine (audio.url), and an optional soundtrack adds a music bed with duck_db from 0 to 20 and fade_out_seconds up to 10. In current code the render takes sound only from the spine and the soundtrack; each clip's own audio is dropped. Building the whole ad around the read is covered in How to make a 30 second advertisement with AI, and an audio-only spot in AI radio commercial generator.
How much does AI voiceover for ads cost?
Text to speech is billed per character of the script at $0.0475 per 1,000 characters; spaces and punctuation count. Each take is billed, so budget for retakes. Synthesized audio longer than 1,200 seconds fails with tts_duration_exceeded and no credit is captured. The script lengths below are example assumptions.
| Example | Requests | Avg. characters | Price |
|---|---|---|---|
| One 30-second read | 1 | 450 | $0.0214 |
| 6, 15 and 30 s cuts of one ad | 3 | 260 | $0.037 |
| 20 variants of a 15-second read | 20 | 230 | $0.2185 |
| 20 variants, 3 takes each | 60 | 230 | $0.6555 |
Can I use an AI voice in a commercial?
What decides it is the terms of the tool you generate with. The Sume Terms of Service say Sume does not claim ownership of generated outputs and that paid plans include commercial use as described on the pricing page. They also say you represent that you have permission to use any person's likeness or voice you submit, which matters for a cloned voice, and that you are responsible for reviewing outputs before commercial use. This is what the terms say, not legal advice; Text to speech commercial use quotes the clauses in full.
Sources
Related posts
More in Use cases
- AI wedding invitation video maker: art, motion, music, text
An AI wedding invitation video is your invitation art animated into short clips, set to music, with names, date and venue burned on as typed text.
- AI whiteboard animation generator: blank board to sketch
AI can imitate whiteboard animation: a blank-board first frame, a finished-sketch last frame, the drawing in between. Words go on as captions.
- Background music for commercials: generate a bed that fits
Background music for a commercial is an instrumental bed under the voiceover that ends on time. How to brief, generate, cut and duck one with AI.
- Can I use AI images on Amazon KDP? What you must disclose
Yes, but KDP requires you to disclose AI-generated images, including cover and interior art. AI-assisted edits of your own work need no disclosure.
Written by Sume