AI voiceover for ads: one brand voice across every cut

Make AI voiceover for ads by pinning one voice and voicing each 6, 15 or 30 second cut as its own text to speech request. Fields, fit and cost.

5 min readSume
All posts

AI voiceover for ads works by picking one voice for the brand and generating each cut's script, whether a 6-, 15- or 30-second version, as its own text to speech request in that voice. You measure the real length of each read, then cut picture to it, and a new variant is a new request rather than a new recording session.

On Sume that request is POST /v1/tts-1.0/generate, billed at $0.0475 per 1,000 characters plus a 5.5% agent fee by default. The fields below come from the TTS schema in the Sume API reference (see API reference), the rate from API pricing, and the mixing step from Timeline 1.0, read on 2026-09-29.

How do I keep the same voice across every ad variant?

Send the same voice selector on every request. There are two: voice.id, which takes a TTS voice UUID or a Voices library id (voi_ plus 32 hex), or the top-level avatar_id / avatar_handle, which uses a ready avatar's voice. If you send both, they must match, or the request fails with 400. A brand voice you clone or describe is made in the Sume app under Assets → Voices, not through the API; see AI voiceover in your own voice.

Keep the other settings fixed across the campaign too, so only the script changes between cuts:

From the TTS 1.0 schema in the Sume API reference, read 2026-09-29.
FieldWhat it does for an ad
voice.id or avatar_idThe brand voice; same value on every cut
languageSet it for every non-English script; omitted means English
generation_config.speedPace multiplier from 0.6 to 1.5
generation_config.volumeVolume multiplier from 0.5 to 2.0
generation_config.emotionA free-form emotion guide, up to 64 characters
output_formatDefaults to MP3 at 44,100 Hz and 128 kbps; WAV is available
timestamps.wordsReturns word timings on the result
transcriptThe script, up to 20,000 characters
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: spring-sale-15s-v1" \
  -d '{
    "transcript": "Spring sale starts Friday. Everything in store, twenty percent off. Visit example.com.",
    "voice": { "mode": "id", "id": "9f3e2b1c-6a4d-4c8e-9f1a-2b3c4d5e6f70" },
    "generation_config": { "speed": 1.1 },
    "output_format": { "container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100 },
    "timestamps": { "words": true },
    "mode": "async"
  }'

How do I make a read fit a 6, 15 or 30 second slot?

Write to the slot, then measure. With timestamps.words set, the completed result carries words[] with start and end seconds, so the last word's end is the spoken length. If a read runs long, cut words first and nudge generation_config.speed second. The estimate-then-measure method is in Text to speech time calculator.

How do I put the voiceover under the ad's picture?

In a Timeline 1.0 render, the voiceover file becomes the audio spine (audio.url), and an optional soundtrack adds a music bed with duck_db from 0 to 20 and fade_out_seconds up to 10. In current code the render takes sound only from the spine and the soundtrack; each clip's own audio is dropped. Building the whole ad around the read is covered in How to make a 30 second advertisement with AI, and an audio-only spot in AI radio commercial generator.

How much does AI voiceover for ads cost?

Text to speech is billed per character of the script at $0.0475 per 1,000 characters; spaces and punctuation count. Each take is billed, so budget for retakes. Synthesized audio longer than 1,200 seconds fails with tts_duration_exceeded and no credit is captured. The script lengths below are example assumptions.

Computed from the Text to speech row on API pricing ($0.0475 per 1,000 characters), read 2026-09-29. Before the 5.5% agent fee.
ExampleRequestsAvg. charactersPrice
One 30-second read1450$0.0214
6, 15 and 30 s cuts of one ad3260$0.037
20 variants of a 15-second read20230$0.2185
20 variants, 3 takes each60230$0.6555

Can I use an AI voice in a commercial?

What decides it is the terms of the tool you generate with. The Sume Terms of Service say Sume does not claim ownership of generated outputs and that paid plans include commercial use as described on the pricing page. They also say you represent that you have permission to use any person's likeness or voice you submit, which matters for a cloned voice, and that you are responsible for reviewing outputs before commercial use. This is what the terms say, not legal advice; Text to speech commercial use quotes the clauses in full.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume