How to make an AI voiceover: from script to audio file

Make an AI voiceover in five steps: write the script, pick a voice, generate the speech, check its length, and download the file. The Sume API way.

5 min readSume
All posts

To make an AI voiceover, write the script to be read aloud, pick a voice, turn the script into speech with a text-to-speech tool, check how long the audio runs, and download the file to place under your video, slides, or ad. With Sume, the speech step is one POST /v1/tts-1.0/generate request carrying the script and a voice. It runs as a job and returns the voiceover as an MP3 by default, or as a WAV if you ask for one.

The request fields come from the TTS schema in the Sume API reference, the OpenAPI document behind the API reference docs, and the polling steps from Jobs and results, read on 2026-09-28. Details called current behavior are read from Sume's code, and the price from the code behind API pricing. Every field of the request is covered in Text to speech API.

How do I write a script for an AI voiceover?

Write it to be heard, not read. The voice speaks the text you send as the transcript, so the script is the only place to control what gets said.

  • Keep sentences short, one idea each, and read the script aloud once before you generate it.
  • Write numbers, abbreviations, and names the way they should sound. Text to speech pronunciation covers fixing a word that comes out wrong.
  • Leave stage directions and notes out: everything in the transcript is spoken.
  • One request takes up to 20,000 characters, and spaces and punctuation count toward usage. Split a longer script, as in text to speech for long text.

Which voice should I use?

Pick a voice whose tone suits the piece and whose language matches the script. Sume's TTS takes the voice in one of two ways:

  • A ready avatar's voice: send its avatar_id or avatar_handle. GET /v1/avatar-1.0/avatars lists your avatars, and any avatar whose voice.status is ready works. This is the voice the API itself lets you discover. In current code an avatar's voice is cloned with its language set to English.
  • A voice id you already hold, sent as voice.id: a voice UUID or a Voices library id (voi_ plus 32 hex characters). In current code, library voices are cloned or designed in the Sume app, not through the API; see AI voiceover in your own voice and AI voice from a text description. Clone only a voice you have permission to use.

How do I generate the voiceover?

Send the script and the voice. Set language for any non-English script; left out, it defaults to English. This request asks for a WAV file and word timings:

  • The default async mode answers at once with status_url and result_url. Poll GET /v1/jobs/{id}/status until terminal is true, and if sume_status is completed, read GET /v1/jobs/{id}/result. A failed or canceled job has no result: /result answers 409 job_not_completed.
  • Completed results expose Sume-hosted audio artifacts. In current code the result's audio_url is the voiceover file on media.sume.com.
  • If your client times out, keep polling rather than sending the request again. A retry with the same Idempotency-Key and the same body returns the original job instead of billing a second one.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: mug-voiceover-001" \
  -d '{
    "transcript": "Meet the new travel mug. It fits every cup holder.",
    "avatar_handle": "product_host",
    "language": "en",
    "generation_config": { "speed": 1.0 },
    "output_format": { "container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100 },
    "timestamps": { "words": true }
  }'

How do I change the pace or check the length?

generation_config.speed is a speed multiplier from 0.6 to 1.5, and generation_config.volume a volume multiplier from 0.5 to 2.0. With timestamps.words: true, the result carries words[], each word's start and end in seconds, and in current code also duration_seconds, the length of the file. Read it before you cut pictures to the voice. Text to speech time calculator shows how to fit a script to a 15-, 30-, or 60-second slot.

What do I do with the voiceover file?

Download it from audio_url and drop it into your editor, or keep working on Sume, where the file is already hosted: add it to a video and keep the video's sound, or mix it with background music into one audio file.

What does an AI voiceover cost, and what are the limits?

Text to speech costs $0.0475 per 1,000 characters, plus a 5.5% agent fee by default, counted on the transcript's characters. Each new take is a new job billed the same way, so test a changed line on its own before you regenerate a long script.

From the TTS 1.0 request schema in the Sume API reference and the code behind API pricing, read 2026-09-28.
LimitValue
Script length1–20,000 characters per request; spaces and punctuation count
Audio lengthOver 1,200 seconds fails with tts_duration_exceeded, and no credits are captured
Speedgeneration_config.speed, 0.6–1.5
Volumegeneration_config.volume, 0.5–2.0
FileMP3 at 44,100 Hz and 128 kbps by default; wav or raw via output_format
DeliveryAn async job with polling or a webhook; not streaming
Price$0.0475 per 1,000 characters

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume