eleven_v4 vs eleven_v4_turbo: model IDs, endpoints, which to pick

ElevenLabs lists eleven_v4 for expressive speech with cloning in 90+ languages and eleven_v4_turbo at about 100 ms median latency. Which fits a video pipeline.

4 min readSume
All posts

Pick eleven_v4 for finished narration where expressiveness and voice cloning matter, and eleven_v4_turbo when latency matters more than the last bit of expression. The ElevenLabs changelog for September 28, 2026, read 2026-10-03, describes v4 as expressive speech with voice cloning in 90+ languages, and v4 Turbo as the real-time variant with a median inference latency of about 100 ms.

What the changelog says

The entry also says which API each model arrived on.

ElevenLabs Sept 28, 2026 changelog (read 2026-10-03)
Model IDDescribed asArrived on
eleven_v4Expressive speech with voice cloning, 90+ languagesText to Dialogue API
eleven_v4_turboReal-time variant, median inference latency about 100 msWebSocket

Which one for video

A video pipeline renders narration ahead of time. Nothing in a finished MP4 benefits from 100 ms latency, because the viewer never waits on the speech. That points to v4 for the voiceover. Turbo earns its place where a person is waiting, such as a live agent or an interactive preview.

Latency is also not the same as turnaround. A pipeline job still includes upload, queueing and your own validation, so the vendor's inference figure is one component of the wait.

Where this lands on Sume

Sume exposes speech through the hosted MCP tts_create tool, which is a paid create that needs an idempotency_key, and routes to sume/auto unless you name a family. The tools doc also lists tts_source_get and tts_source_verify_spine, which check selected TTS jobs against an accepted script. The docs reviewed for this post do not list ElevenLabs model ids, so confirm which voices and models your workspace can use before you plan around a specific id.

After narration lines exist as separate files, timeline audio can concat 1-20 ordered Sume-hosted parts into one gapless file with sample-domain joins and no re-synthesis, at $0.01 flat per job.

Testing both models

Render the same 30 seconds of script with each model and listen on the device your audience uses, usually a phone speaker. Check names, numbers and the first and last words of each line, where synthesis most often slips. Save the model id with each file so you can rerun the comparison after a model update.

If the two sound equivalent on your script, choose on latency and price, not on the label.

A decision rule

Ask whether a human waits on the audio. If yes, use the low-latency model. If no, use the expressive one and pre-render. Keep the model id in your job metadata so you can reproduce a take later.

Sources

Related posts

More in Models

All Models posts

Written by Sume