Make an AI voice sound the same across videos: read the job record
A completed Sume TTS job records the engine, voice, language, output format and settings it used. Reuse them so the next line matches. Why it matters after v4.

To make a voice sound the same across a series, reuse the settings of a take you liked, and Sume gives them to you: a completed text-to-speech job records model_id, voice, language, output_format, generation_config and speed, each set to null when the request did not send it. Read those from the job and send them again on the next line. Model changes make this more important: ElevenLabs launched v4 on 2026-09-28 with stackable inline tags, and new expressive controls are exactly what shifts a voice between episodes.
This is a habit, not a guarantee; check the audio.
Which settings does Sume record?
The Jobs and results docs list them. Settings you did not send are recorded as null, which tells you the engine used its own default, so do not try to copy a null.
| Recorded field | What it holds | Request field to reuse it |
|---|---|---|
| model_id | The engine that made the audio | model, on the TTS Router |
| voice | mode id plus the voice id | avatar_handle, avatar_id or voice.id |
| language | The language of the transcript | language |
| output_format | Container, sample rate, encoding | output_format |
| generation_config | Volume, speed, emotion, or null | generation_config |
| speed | The older slow, normal, fast enum, or null | Prefer generation_config.speed |
How do I reuse a take's settings?
Fetch the result of the take you liked, copy the recorded non-null fields into your next request, and add the new transcript. Use the TTS Router if you need to pin the engine, because the main TTS route has no engine picker and always uses the engine Sume currently chooses. Within the router, sonic-latest is an alias that moves, so pin an explicit id such as sonic-3.6 for a series.
Speed has a range of 0.6 to 1.5 and volume 0.5 to 2.0 in generation_config; emotion is a free-text guide up to 64 characters.
What else drifts between episodes?
The script's punctuation and the voice's language. A voice set up for one language can warn or refuse when you request another; Sume returns a language-mismatch response before any charge unless you confirm it. And a pronunciation fix only holds if you keep the same pronunciation dictionary id. Store those with the rest. See Jobs and results and the API reference.
What should I do?
Save the job id of your reference take in the series bible. Before each episode, render one test line with the recorded settings and compare it to the reference before you render the full script.
Sources
Related posts
More in Developers
- Mastra 1.69 Inngest retries option: durable agents and Sume jobs
Mastra core 1.69.0 adds a retries option for Inngest workflows and durable agents. Pair it with a Sume Idempotency-Key so a retried step never pays twice.
- Check a Meta ad batch for duplicates before upload with Python
A short Python check that flags ads in a batch sharing the same hook and visual before you upload. Useful now that Meta rates creative diversity.
- Midjourney collaborative generation: who can see jobs on Sume?
Midjourney plans ways to generate together. On Sume, an API key reads only the jobs its own member created, so sharing takes a plan. Here are the rules.
- Midjourney edit history: keep an edit lineage with Sume job metadata
Midjourney added edit-history navigation and plans persistent edit history. For the Sume image API, build your own lineage with metadata and the jobs list.
Written by Sume