eleven_v4 vs eleven_v4_turbo: model IDs, endpoints, which to pick
ElevenLabs lists eleven_v4 for expressive speech with cloning in 90+ languages and eleven_v4_turbo at about 100 ms median latency. Which fits a video pipeline.

Pick eleven_v4 for finished narration where expressiveness and voice cloning matter, and eleven_v4_turbo when latency matters more than the last bit of expression. The ElevenLabs changelog for September 28, 2026, read 2026-10-03, describes v4 as expressive speech with voice cloning in 90+ languages, and v4 Turbo as the real-time variant with a median inference latency of about 100 ms.
What the changelog says
The entry also says which API each model arrived on.
| Model ID | Described as | Arrived on |
|---|---|---|
eleven_v4 | Expressive speech with voice cloning, 90+ languages | Text to Dialogue API |
eleven_v4_turbo | Real-time variant, median inference latency about 100 ms | WebSocket |
Which one for video
A video pipeline renders narration ahead of time. Nothing in a finished MP4 benefits from 100 ms latency, because the viewer never waits on the speech. That points to v4 for the voiceover. Turbo earns its place where a person is waiting, such as a live agent or an interactive preview.
Latency is also not the same as turnaround. A pipeline job still includes upload, queueing and your own validation, so the vendor's inference figure is one component of the wait.
Where this lands on Sume
Sume exposes speech through the hosted MCP tts_create tool, which is a paid create that needs an idempotency_key, and routes to sume/auto unless you name a family. The tools doc also lists tts_source_get and tts_source_verify_spine, which check selected TTS jobs against an accepted script. The docs reviewed for this post do not list ElevenLabs model ids, so confirm which voices and models your workspace can use before you plan around a specific id.
After narration lines exist as separate files, timeline audio can concat 1-20 ordered Sume-hosted parts into one gapless file with sample-domain joins and no re-synthesis, at $0.01 flat per job.
Testing both models
Render the same 30 seconds of script with each model and listen on the device your audience uses, usually a phone speaker. Check names, numbers and the first and last words of each line, where synthesis most often slips. Save the model id with each file so you can rerun the comparison after a model update.
If the two sound equivalent on your script, choose on latency and price, not on the label.
A decision rule
Ask whether a human waits on the audio. If yes, use the low-latency model. If no, use the expressive one and pre-render. Keep the model id in your job metadata so you can reproduce a take later.
Sources
Related posts
More in Models
- ElevenLabs languages: Flash v2.5 has 32, Multilingual v2 29, v4 90+
ElevenLabs lists 32 languages for Flash v2.5, 29 for Multilingual v2 and 90+ for v4 and v4 Turbo. Check your markets against the model, then log it per job.
- Gemini 3.8 Flash TTS tops Hume's VoiceEQ board: run your own test
Hume's blog lists Gemini 3.8 Flash TTS atop its Real-World VoiceEQ board. Why a vendor-run board is only a lead, and how to run a blind A/B on your script.
- Gemini Omni Flash GA: extension and interpolation vs Sume's inputs
Gemini Omni Flash is generally available with extension and interpolation between images. What Sume's gemini-omni-flash-1.1 catalog entry documents instead.
- Gemini Omni Flash went GA Aug 27: five checks after a preview ends
Omni Flash entered public preview Jun 30 and went GA Aug 27 with extension and 360p-4K. Re-test these five limits before you reuse numbers from preview runs.
Written by Sume