How to make an AI voiceover: from script to audio file
Make an AI voiceover in five steps: write the script, pick a voice, generate the speech, check its length, and download the file. The Sume API way.

To make an AI voiceover, write the script to be read aloud, pick a voice, turn the script into speech with a text-to-speech tool, check how long the audio runs, and download the file to place under your video, slides, or ad. With Sume, the speech step is one POST /v1/tts-1.0/generate request carrying the script and a voice. It runs as a job and returns the voiceover as an MP3 by default, or as a WAV if you ask for one.
The request fields come from the TTS schema in the Sume API reference, the OpenAPI document behind the API reference docs, and the polling steps from Jobs and results, read on 2026-09-28. Details called current behavior are read from Sume's code, and the price from the code behind API pricing. Every field of the request is covered in Text to speech API.
How do I write a script for an AI voiceover?
Write it to be heard, not read. The voice speaks the text you send as the transcript, so the script is the only place to control what gets said.
- Keep sentences short, one idea each, and read the script aloud once before you generate it.
- Write numbers, abbreviations, and names the way they should sound. Text to speech pronunciation covers fixing a word that comes out wrong.
- Leave stage directions and notes out: everything in the transcript is spoken.
- One request takes up to 20,000 characters, and spaces and punctuation count toward usage. Split a longer script, as in text to speech for long text.
Which voice should I use?
Pick a voice whose tone suits the piece and whose language matches the script. Sume's TTS takes the voice in one of two ways:
- A ready avatar's voice: send its
avatar_idoravatar_handle.GET /v1/avatar-1.0/avatarslists your avatars, and any avatar whosevoice.statusisreadyworks. This is the voice the API itself lets you discover. In current code an avatar's voice is cloned with its language set to English. - A voice id you already hold, sent as
voice.id: a voice UUID or a Voices library id (voi_plus 32 hex characters). In current code, library voices are cloned or designed in the Sume app, not through the API; see AI voiceover in your own voice and AI voice from a text description. Clone only a voice you have permission to use.
How do I generate the voiceover?
Send the script and the voice. Set language for any non-English script; left out, it defaults to English. This request asks for a WAV file and word timings:
- The default
asyncmode answers at once withstatus_urlandresult_url. PollGET /v1/jobs/{id}/statusuntilterminalis true, and ifsume_statusiscompleted, readGET /v1/jobs/{id}/result. A failed or canceled job has no result:/resultanswers409 job_not_completed. - Completed results expose Sume-hosted audio artifacts. In current code the result's
audio_urlis the voiceover file onmedia.sume.com. - If your client times out, keep polling rather than sending the request again. A retry with the same
Idempotency-Keyand the same body returns the original job instead of billing a second one.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: mug-voiceover-001" \
-d '{
"transcript": "Meet the new travel mug. It fits every cup holder.",
"avatar_handle": "product_host",
"language": "en",
"generation_config": { "speed": 1.0 },
"output_format": { "container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100 },
"timestamps": { "words": true }
}'How do I change the pace or check the length?
generation_config.speed is a speed multiplier from 0.6 to 1.5, and generation_config.volume a volume multiplier from 0.5 to 2.0. With timestamps.words: true, the result carries words[], each word's start and end in seconds, and in current code also duration_seconds, the length of the file. Read it before you cut pictures to the voice. Text to speech time calculator shows how to fit a script to a 15-, 30-, or 60-second slot.
What do I do with the voiceover file?
Download it from audio_url and drop it into your editor, or keep working on Sume, where the file is already hosted: add it to a video and keep the video's sound, or mix it with background music into one audio file.
What does an AI voiceover cost, and what are the limits?
Text to speech costs $0.0475 per 1,000 characters, plus a 5.5% agent fee by default, counted on the transcript's characters. Each new take is a new job billed the same way, so test a changed line on its own before you regenerate a long script.
| Limit | Value |
|---|---|
| Script length | 1–20,000 characters per request; spaces and punctuation count |
| Audio length | Over 1,200 seconds fails with tts_duration_exceeded, and no credits are captured |
| Speed | generation_config.speed, 0.6–1.5 |
| Volume | generation_config.volume, 0.5–2.0 |
| File | MP3 at 44,100 Hz and 128 kbps by default; wav or raw via output_format |
| Delivery | An async job with polling or a webhook; not streaming |
| Price | $0.0475 per 1,000 characters |
Sources
Related posts
More in Use cases
- How to make an unboxing video with AI, step by step
An unboxing video shows hands opening the package and revealing the product. How to make one with AI from a box photo and a product photo.
- How to create meeting minutes from an audio recording
Create meeting minutes from a recording in two steps: transcribe the audio, then have a language model draft the summary, decisions, and action items.
- Create motivational videos with AI: voice, shots, and music
Create motivational videos with AI: a slow spoken quote over cinematic shots, a music bed that builds and ducks under the voice, and big captions.
- Press release video: an announcement read by an AI avatar
A press release video reads the headline, key facts, and a quote in about a minute, captioned from the release text. How to make one with an AI avatar.
Written by Sume