AI voiceover preflight checklist: 12 checks before you render

A voiceover is cheap to redo and costly to find wrong after a render. Twelve checks tied to Sume TTS limits and error codes, from the 20,000-character cap down.

4 min readSume
All posts

Before you spend $0.10 a minute on a render, run these twelve checks on the voiceover. Each one maps to a limit or error code in the Sume TTS contract, so none of them is guesswork (API reference). A failed check costs a few cents at $0.0475 per 1,000 characters to fix, where a bad voiceover found after the render costs the render too.

The checklist

Twelve preflight checks and the Sume field or error behind each (read 2026-10-04)
CheckLimit or codeWhat to do
Script length20,000 characters per requestSplit by chapter
Audio length1,200 s, else tts_duration_exceededSplit, then join with timeline audio
Language setlanguage for non-English textSend it every time
Voice fits languagetts_voice_language_warningListen, then confirm or change voice
Voice selectoravatar and voice.id must matchSend one, not both
Idempotency keyRequired on paid writesOne key per line and version
Cost previewdry_run, max_spend_usd on tts_createPreview before a batch
File typewav joins cleanly, mp3 adds priming paddingUse wav for joins and renders
Timingtimestamps.wordsRequest it if captions follow the audio
Round tripSTT of the takeTranscribe it and diff against the script
Render lengthaudio.duration_seconds 1 to 1,800Round up from the TTS duration
Fallback pollGET /v1/jobs/{id}/statusKeep it beside any webhook

The two checks people skip

First, the round trip. Run the finished take through Sume STT at $0.01 per audio minute and compare it with the script. A skipped or mangled word shows up as a diff, and the round-trip guide has a Python script for it.

Second, the length check. The render takes audio.duration_seconds as an integer from 1 to 1,800, and it bills per output minute, rounded up. Use the TTS duration_seconds and round up before you submit (timeline docs).

Why a list helps more as models get faster

A fast model makes it tempting to skip review. Microsoft quotes 150 ms for 45 seconds of audio on MAI-Voice-2.1-Flash (Microsoft AI, read 2026-10-04). Speed of generation does not reduce the time it takes a person to hear a voiceover once. Keep the list next to your scripts and tick it per episode.

The hosted MCP gates, idempotency_key, dry_run and max_spend_usd, are documented in the MCP tools page. They cover the cost items on this list when an agent drives the calls.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume