Eleven v4 and Sume: a step-by-step feature map for voiceover work

Eleven v4 is not a model id on Sume. Here is what v4 does per ElevenLabs, and which Sume endpoint covers each step of a voiceover pipeline today.

4 min readSume
All posts

Eleven v4 is not a model id you can call on Sume today. The Sume TTS Router catalog is Cartesia Sonic only, so the honest answer is a feature map: what ElevenLabs says v4 does, and which Sume step covers the same job with Sonic.

Use this when someone in your team asks for v4 and you need to say what is possible without it.

What ElevenLabs says about v4

The Eleven v4 announcement (published 2026-09-28) describes inline tags such as [laughs] and [light rain], IPA phoneme control, 90+ languages, and Instant Voice Clone from about ten seconds of audio. The models doc lists eleven_v4 with 90+ languages and a 10,000-character limit, and eleven_v4_turbo at about 100 ms.

The feature map

Sume's relevant surface is the TTS Router and TTS 1.0 contract. Where the table says no equivalent documented, Sume does not offer it as a documented feature.

Eleven v4 features and the Sume step that covers the job (read 2026-10-03)
JobElevenLabs v4Sume today
Pick a modeleleven_v4, eleven_v4_turbosonic-3.6, sonic-3.5, sonic-3, sonic-latest, sonic-preview
Characters per request10,000 for eleven_v420,000 transcript limit
Languages90+Sonic languages, set with a BCP-47 language code
Delivery controlInline tagsgeneration_config: volume, speed, emotion
PronunciationIPA phonemespronunciation_dict_id field
Word timingsNot covered heretimestamps.words true
Sound design in the text[light rain] style tagsNo equivalent documented

Where the gap is real

Inline sound-effect tags and the 90+ language count are not things Sume documents for its catalog. If your project depends on them, use ElevenLabs directly and bring the audio into Sume as a file URL for the timeline. If your need is narration with predictable timing, captions and a clean video pipeline, the Sonic route is sufficient.

  • Need sound effects in speech: do that outside Sume.
  • Need word timings for captions: Sume returns them.
  • Need long narrations: split under the 20,000 character and 1200 second limits.
  • Need cloning: see what Sume offers for voice cloning.

Decision shortcut

If the deliverable is a video, the extra steps after TTS (captions, timeline, music) matter more than the last few points of voice preference. Run the blind test from the leaderboard post on your script and decide with ears, not badges.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume