Make an AI voice sound the same across videos: read the job record

A completed Sume TTS job records the engine, voice, language, output format and settings it used. Reuse them so the next line matches. Why it matters after v4.

4 min readSume
All posts

To make a voice sound the same across a series, reuse the settings of a take you liked, and Sume gives them to you: a completed text-to-speech job records model_id, voice, language, output_format, generation_config and speed, each set to null when the request did not send it. Read those from the job and send them again on the next line. Model changes make this more important: ElevenLabs launched v4 on 2026-09-28 with stackable inline tags, and new expressive controls are exactly what shifts a voice between episodes.

This is a habit, not a guarantee; check the audio.

Which settings does Sume record?

The Jobs and results docs list them. Settings you did not send are recorded as null, which tells you the engine used its own default, so do not try to copy a null.

Fields on a completed Sume TTS job, from the Jobs and results docs and OpenAPI schema, read 2026-10-02.
Recorded fieldWhat it holdsRequest field to reuse it
model_idThe engine that made the audiomodel, on the TTS Router
voicemode id plus the voice idavatar_handle, avatar_id or voice.id
languageThe language of the transcriptlanguage
output_formatContainer, sample rate, encodingoutput_format
generation_configVolume, speed, emotion, or nullgeneration_config
speedThe older slow, normal, fast enum, or nullPrefer generation_config.speed

How do I reuse a take's settings?

Fetch the result of the take you liked, copy the recorded non-null fields into your next request, and add the new transcript. Use the TTS Router if you need to pin the engine, because the main TTS route has no engine picker and always uses the engine Sume currently chooses. Within the router, sonic-latest is an alias that moves, so pin an explicit id such as sonic-3.6 for a series.

Speed has a range of 0.6 to 1.5 and volume 0.5 to 2.0 in generation_config; emotion is a free-text guide up to 64 characters.

What else drifts between episodes?

The script's punctuation and the voice's language. A voice set up for one language can warn or refuse when you request another; Sume returns a language-mismatch response before any charge unless you confirm it. And a pronunciation fix only holds if you keep the same pronunciation dictionary id. Store those with the rest. See Jobs and results and the API reference.

What should I do?

Save the job id of your reference take in the series bible. Before each episode, render one test line with the recorded settings and compare it to the reference before you render the full script.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume