Freeze a TTS preset so every episode sounds the same

Keep one JSON preset for a recurring voiceover: avatar handle, language, speed 0.6-1.5, volume 0.5-2, output format. Sume also records these values on each job.

4 min readSume
All posts

To make every episode of a recurring voiceover sound the same, store one JSON preset and send it unchanged with each script. On Sume that preset is the voice selector (avatar_handle, avatar_id or voice.id), language, generation_config (speed 0.6 to 1.5, volume 0.5 to 2, an optional emotion string of up to 64 characters) and output_format. A completed TTS job records the model, voice, language, output format and generation config it used, so you can read the values back from the job to match the next line (API reference).

What belongs in the preset

Only fields that change the sound or the file. Leave the transcript out, since that changes per episode.

  • Voice: the avatar handle or voice id. If you send both an avatar and voice.id, they must match.
  • Language: required for any non-English text. A mismatch between voice and language raises tts_voice_language_warning unless you send confirm_language_mismatch.
  • Speed and volume: numeric, in the ranges above. The older speed enum (slow, normal, fast) is deprecated, so use generation_config.speed.
  • Output format: container, sample rate and, for mp3, bit rate.
{
  "avatar_handle": "your-voice-handle",
  "language": "en",
  "generation_config": {"speed": 0.95, "volume": 1.0},
  "output_format": {"container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100}
}

The model is a separate question

TTS 1.0 has no engine picker, and sending model returns a 400. If you need to hold the engine steady across a season, use POST /v1/tts-router/generate with an explicit model such as sonic-3.6. sonic-latest is an alias for the newest model and can move, which is why the pin guide recommends naming the version.

Why this matters now

New voices keep shipping. Microsoft's MAI-Voice-2.1 is described as one voice that works across 23 languages and 26 locales (Microsoft AI, read 2026-10-04). A faster or newer model is a reason to test, and a frozen preset is how you compare fairly, by changing one field at a time.

A habit that works

Version the preset file next to your scripts. When an episode sounds off, read generation_config from its job and diff it against the file before blaming the voice. A 200-character test line costs under a cent at $0.0475 per 1,000 characters, so run one before each new season.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume