AI voiceover preflight checklist: 12 checks before you render
A voiceover is cheap to redo and costly to find wrong after a render. Twelve checks tied to Sume TTS limits and error codes, from the 20,000-character cap down.

Before you spend $0.10 a minute on a render, run these twelve checks on the voiceover. Each one maps to a limit or error code in the Sume TTS contract, so none of them is guesswork (API reference). A failed check costs a few cents at $0.0475 per 1,000 characters to fix, where a bad voiceover found after the render costs the render too.
The checklist
| Check | Limit or code | What to do |
|---|---|---|
| Script length | 20,000 characters per request | Split by chapter |
| Audio length | 1,200 s, else tts_duration_exceeded | Split, then join with timeline audio |
| Language set | language for non-English text | Send it every time |
| Voice fits language | tts_voice_language_warning | Listen, then confirm or change voice |
| Voice selector | avatar and voice.id must match | Send one, not both |
| Idempotency key | Required on paid writes | One key per line and version |
| Cost preview | dry_run, max_spend_usd on tts_create | Preview before a batch |
| File type | wav joins cleanly, mp3 adds priming padding | Use wav for joins and renders |
| Timing | timestamps.words | Request it if captions follow the audio |
| Round trip | STT of the take | Transcribe it and diff against the script |
| Render length | audio.duration_seconds 1 to 1,800 | Round up from the TTS duration |
| Fallback poll | GET /v1/jobs/{id}/status | Keep it beside any webhook |
The two checks people skip
First, the round trip. Run the finished take through Sume STT at $0.01 per audio minute and compare it with the script. A skipped or mangled word shows up as a diff, and the round-trip guide has a Python script for it.
Second, the length check. The render takes audio.duration_seconds as an integer from 1 to 1,800, and it bills per output minute, rounded up. Use the TTS duration_seconds and round up before you submit (timeline docs).
Why a list helps more as models get faster
A fast model makes it tempting to skip review. Microsoft quotes 150 ms for 45 seconds of audio on MAI-Voice-2.1-Flash (Microsoft AI, read 2026-10-04). Speed of generation does not reduce the time it takes a person to hear a voiceover once. Keep the list next to your scripts and tick it per episode.
The hosted MCP gates, idempotency_key, dry_run and max_spend_usd, are documented in the MCP tools page. They cover the cost items on this list when an agent drives the calls.
Sources
Related posts
More in Developers
- Profit per SKU across channels: add AI media cost from Sume
Amazon lists cross-channel profitability as upcoming. Add what your product images and clips cost per SKU by reading Sume's usage ledger by job id.
- Animate a still by API: first-frame jobs on Sume
fal lists FLUX 3 as animating one still into video. On Sume you send the still as a first frame to a video model; this Python script submits and polls.
- Ask for the aspect ratio first: an MCP input_required round trip
MCP multi round-trip requests let a tool answer input_required to ask for an aspect ratio or spend approval before a render, with state in requestState.
- Attach a terminal to a run: ant sessions connect vs Sume jobs
The ant CLI can attach to a Managed Agents session. For Sume generation jobs, use sume jobs watch and MCP jobs_wait instead, and never resubmit a paid job.
Written by Sume