Freeze a TTS preset so every episode sounds the same
Keep one JSON preset for a recurring voiceover: avatar handle, language, speed 0.6-1.5, volume 0.5-2, output format. Sume also records these values on each job.

To make every episode of a recurring voiceover sound the same, store one JSON preset and send it unchanged with each script. On Sume that preset is the voice selector (avatar_handle, avatar_id or voice.id), language, generation_config (speed 0.6 to 1.5, volume 0.5 to 2, an optional emotion string of up to 64 characters) and output_format. A completed TTS job records the model, voice, language, output format and generation config it used, so you can read the values back from the job to match the next line (API reference).
What belongs in the preset
Only fields that change the sound or the file. Leave the transcript out, since that changes per episode.
- Voice: the avatar handle or voice id. If you send both an avatar and
voice.id, they must match. - Language: required for any non-English text. A mismatch between voice and language raises
tts_voice_language_warningunless you sendconfirm_language_mismatch. - Speed and volume: numeric, in the ranges above. The older
speedenum (slow,normal,fast) is deprecated, so usegeneration_config.speed. - Output format: container, sample rate and, for mp3, bit rate.
{
"avatar_handle": "your-voice-handle",
"language": "en",
"generation_config": {"speed": 0.95, "volume": 1.0},
"output_format": {"container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100}
}The model is a separate question
TTS 1.0 has no engine picker, and sending model returns a 400. If you need to hold the engine steady across a season, use POST /v1/tts-router/generate with an explicit model such as sonic-3.6. sonic-latest is an alias for the newest model and can move, which is why the pin guide recommends naming the version.
Why this matters now
New voices keep shipping. Microsoft's MAI-Voice-2.1 is described as one voice that works across 23 languages and 26 locales (Microsoft AI, read 2026-10-04). A faster or newer model is a reason to test, and a frozen preset is how you compare fairly, by changing one field at a time.
A habit that works
Version the preset file next to your scripts. When an episode sounds off, read generation_config from its job and diff it against the file before blaming the voice. A 200-character test line costs under a cent at $0.0475 per 1,000 characters, so run one before each new season.
Sources
Related posts
More in Developers
- Gemini 3.8 Live audio: wrap 24 kHz PCM in WAV, resample to 16 kHz
Gemini 3.8 Live takes 16-bit 16 kHz PCM in and returns 24 kHz out. A Python WAV wrapper, an ffmpeg resample command, and the Sume detach settings that match.
- New model id on a Sume Format run: probe for a 400 before launch day
When Gemini 4 Argon or any new model opens up, test whether a Format run accepts its id. A 400 invalid_request means it is not in the catalog yet.
- Gemini video understanding 88% fewer tokens vs Sume Video inspect
Gemini reports up to 88% fewer tokens on long video. Sume Video inspect and Reference ingest take another route: stills, transcript and a manifest.
- Gemini API paid vs unpaid data use: Omni prompts and uploaded clips
Gemini API terms: unpaid content may improve Google products and reach human reviewers; paid content does not. What that means for Omni edit uploads.
Written by Sume