Keep a series voice consistent: pin model, voice, speed and volume
A series sounds the same only if every episode sends the same TTS settings. Keep one profile in code, pin a model id, and send it with each Sume request.

To keep a voice consistent across a series, store one profile object in your code and send all of it on every Sume TTS request: the model id, the voice (an avatar_handle or voice.id), the language, and the generation_config with speed, volume and emotion. Pin the model to a dated id such as sonic-3.6, not the sonic-latest alias. The router docs say the alias resolves to sonic-3.6 today, but an alias is a pointer that the docs do not promise to hold still, while the dated id is a fixed choice. The profile below is plain Python and is shared by every episode script.
This matters now because series are the format platforms are pushing. YouTube Shorts series, with seasons, episodes and sequential playback, began rolling out from 2026-09-23 (Orthotropy, read 2026-10-06). A viewer who plays episode 1 then episode 2 in a row hears every difference in the voice.
One profile, every episode
import json, os, time, urllib.request as u
KEY = os.environ["SUME_API_KEY"]
def call(url, body=None, key=None):
h = {"Authorization": "Bearer " + KEY, "Content-Type": "application/json"}
if key: h["Idempotency-Key"] = key
data = json.dumps(body).encode() if body else None
return json.load(u.urlopen(u.Request(url, data=data, headers=h)))["data"]
PROFILE = {
"model": "sonic-3.6",
"language": "en",
"avatar_handle": os.environ["SUME_AVATAR_HANDLE"],
"generation_config": {"speed": 1.0, "volume": 1.0},
"output_format": {"container": "wav"},
}
def voice_episode(n, transcript):
body = dict(PROFILE, transcript=transcript)
job = call("https://api.sume.com/v1/tts-router/generate", body, f"series-a-ep{n}-v1")
while not job["terminal"]:
time.sleep(job.get("next_poll_after_seconds") or 2)
job = call(job["status_url"])
arts = call(job["result_url"])["result"]["artifacts"]
return next(a["url"] for a in arts if a["type"] == "audio")
episodes = {1: "Episode one. The setup.", 2: "Episode two. The twist."}
for n, text in episodes.items():
print(n, voice_episode(n, text))What carries the sound
Three fields carry most of the consistency. The model: each Sume model is a different engine, so a switch changes the sound even on the same voice id, which is why a pinned id belongs in the profile. The voice: use the same handle every time, and check that its status is ready. The settings: speed runs 0.6 to 1.5 and volume 0.5 to 2, and a drift of a tenth in either is audible when clips play in order. The key includes the episode number and a version, so a re-render of episode 2 does not return episode 1's file.
Guard the profile
Keep the profile in one module and import it everywhere, so no script can send a stray value. Add a test that renders the same short line with the stored profile and compares its duration with the previous run: a changed duration is the cheapest alarm that a setting moved. A line of 30 characters costs a fraction of a cent, so run it before each episode batch, not once a quarter.
Changing the profile deliberately
Change the profile on purpose, with a version note. When a new model arrives, render the same test line with the old profile and the new one, listen back to back, and move the whole series only if you want the change. An alias suits a one-off job where you do not care which Sonic answers; a series is the opposite case. The listed price is the same across these models: Sume TTS is $0.0475 per 1,000 characters (pricing, read 2026-10-06), so the choice is about sound and not about cost. If you later move a series to a different model, the budget does not change; only the sound does.
| Field | Pin it to | Why |
|---|---|---|
| model | sonic-3.6 (not sonic-latest) | An alias is a pointer, not a pin |
| avatar_handle or voice.id | One value | Same speaker every episode |
| language | One code | Avoids the 409 language guard |
| generation_config | speed, volume, emotion | Audible drift otherwise |
| output_format | wav | One format into the editor |
Pair it with caption pinning
The captions deserve the same treatment: see same caption look on every episode for pinning style and design. Sume's job envelope, used by the script above, is documented in Jobs and results. For the choice between a pinned id and the alias, read sonic-latest vs sonic-3.6.
Sources
Related posts
More in Developers
- Same volume every Shorts episode: gain_db, duck_db, one music bed
Sume does not measure loudness. To keep a series consistent, pin audio.gain_db, soundtrack gain_db and duck_db in one function with one bed. Python plan loop.
- Ktor webhook route for an AI video job: receiveText and HMAC SHA-256
A Ktor route reads the raw body with call.receiveText(), checks Sume's sume-v1 HMAC over timestamp.body in Kotlin and returns 401 when the secret is empty.
- Label speakers without diarization: transcribe each mic track, merge
Sume STT has no speaker labels. If you record each speaker on a separate track, transcribe both tracks and merge the segments by start time. Python script.
- Laravel queued job for an AI video API: delay, redispatch, poll
A Laravel ShouldQueue job reads a Sume video job once and redispatches itself with ->delay() from next_poll_after_seconds, so no worker sleeps during a render.
Written by Sume