ElevenLabs made eleven_v4_turbo a default: pin your TTS engine
ElevenLabs set eleven_v4_turbo as a default on Oct 5. Sume carries Sonic engines, not Eleven; here is how to pin an engine id so audio does not drift.

A provider changing a default model changes your audio without a code change; the fix is to name the engine in every request. ElevenLabs' changelog for October 5, 2026 says it added eleven_v4 and eleven_v4_turbo and changed the conversational default from eleven_flash_v2 to eleven_v4_turbo. Sume does not list any ElevenLabs model, so the Sume side of this story is about its own engine ids.
Sources: the ElevenLabs changelog entry (read 2026-10-10) and Sume's API reference.
What changed at ElevenLabs?
The changelog entry is short. It names two new text-to-speech models and one default change. It says nothing there about price or voice compatibility, so this post draws no conclusion about either.
| Change | Detail |
|---|---|
| Added model | eleven_v4 |
| Added model | eleven_v4_turbo |
| Default changed | Conversational default from eleven_flash_v2 to eleven_v4_turbo |
How does Sume avoid silent default changes?
Sume TTS 1.0 has no engine picker; it rejects model and model_id, so it is the route to use when you do not care which Sonic release speaks. When you do care, call the TTS Router at POST /v1/tts-router/generate, where model is required. The accepted ids are sonic-3.6, sonic-3.5, sonic-3, sonic-latest and sonic-preview.
sonic-latest is the moving one: the schema calls it an alias for sonic-3.6, and an alias can move when a newer stable release ships. A fixed id like sonic-3.6 does not carry that risk. sonic-preview is a beta id and is incompatible with pro voice clones, which returns voice_model_mismatch.
- Want repeatable narration: send
sonic-3.6(or whichever fixed id you tested). - Want the newest stable one automatically: send
sonic-latest, and log what you got. - Never ship a production pipeline on
sonic-preview.
What to log with every job
Treat the engine id as data, not config you forget. Store the request you sent, including model, voice selector, language and any generation_config, next to the result audio. When a listener says the voice changed in October, you can then tell a changed engine id from a changed script. The price is the same across the router ids, $0.0475 per 1,000 characters, so pinning costs nothing extra.
| Id | Note |
|---|---|
sonic-3.6 | Fixed release |
sonic-3.5 | Fixed release |
sonic-3 | Fixed release |
sonic-latest | Alias for sonic-3.6 |
sonic-preview | Beta; not for pro voice clones |
Is there a way to see the current list?
Yes: GET /v1/tts-router/models returns the live catalog, so a deploy script can check that the id it pins still exists. If you must have an Eleven model, use ElevenLabs directly and bring the finished file into a Sume timeline or caption job as a public HTTPS audio source.
What does a pinned request look like?
A minimal TTS Router call needs the engine, a transcript and a voice. Set model to a fixed id, transcript to your text, and either an avatar or voice.id. Add language for anything not English. The response is a job with status, result and cancel URLs; poll the status URL until the result is ready, then download the audio. Nothing in the request body asks for the newest engine unless you wrote sonic-latest.
For a team, the practical rule is one constant in code, such as TTS_ENGINE = "sonic-3.6", referenced by every call, plus a deploy-time check against GET /v1/tts-router/models. When you decide to move to a newer release, change the constant, regenerate a short test line, listen, and only then roll it out to scheduled jobs.
- One engine constant per brand voice.
- A test line per language before any rollout.
- The changelog of your vendor is a trigger for a listening test, not for a code change.
Sources
Related posts
More in Developers
- ElevenLabs dubbing 180-minute app limit vs 3 GB API and Sume detach
ElevenLabs dubs up to 180 minutes in the app or 3 GB via API. Sume detach takes 1,800 s of source and 900 s of output per job, so long videos need ranges.
- ElevenLabs dubbing output: mono v1, stereo v2, and Sume channels
ElevenLabs dubbing v1 returns mono and v2 at most stereo. On Sume, detach keeps source channels or mono, and a concat of mixed layouts fails by design.
- Scribe accepts MP4 and MOV; Sume STT needs an audio URL: detach first
ElevenLabs Scribe takes video files such as MP4 and MOV directly. Sume STT wants a public HTTPS audio_url, so a video goes through audio detach first at $0.01.
- ElevenLabs 8 kHz mu-law and A-law vs Sume TTS pcm_mulaw, pcm_alaw
For phone-line audio, ElevenLabs lists 8 kHz mu-law and A-law. Sume TTS 1.0 offers pcm_mulaw and pcm_alaw in a raw container at 8000 Hz.
Written by Sume