Eleven v4 drops the native accent; Sume keeps a language tag per voice
ElevenLabs' docs say v4 gives fluent target-language speech, not a preserved accent. Sume tags each voice with a language and checks for mismatch.

ElevenLabs' Eleven v4 docs say v4 produces fluent speech in the target language rather than preserving the speaker's native accent (read 2026-10-03). So a cloned voice speaking another language will sound like a native speaker of it, not like the original speaker with an accent. Sume works differently on the surface: each voice carries one language tag, you send a language code with the transcript, and a mismatch between the two returns a warning before any charge.
Side by side
The ElevenLabs statements are from its model docs, and the Sume ones from the MCP tools and gates page.
| Question | Eleven v4 docs | Sume tts_create |
|---|---|---|
| Cross-language output | Fluent target-language speech, no preserved native accent | Voice has one language tag; send matching language |
| Language omitted | Not stated on the page | Provider assumes English |
| Voice and language disagree | Not stated on the page | Warning, no job and no charge until confirmed |
Which one you want
If a spokesperson should keep their accent across languages, neither page promises that. The ElevenLabs page says the opposite. For a multilingual channel, the safer route on Sume is one voice per language, each tagged, so the voice and the transcript agree. Listen to a sample before a batch in either service.
Sources
Related posts
More in Comparisons
- Eleven v4 voice clone: 10 seconds or 1-2 minutes? Pages differ
ElevenLabs' launch post says an instant clone needs 10 seconds of audio; its docs page says one to two minutes. What Sume's Voices clone asks for instead.
- ElevenLabs v4 Turbo vs Flash vs v3: price per 1,000 characters
ElevenLabs API list per 1K characters: v4 Turbo $0.011 until Oct 12 (then $0.04), Flash/Turbo $0.04, v3 Multilingual $0.08, v4 $0.022 (then $0.08).
- Face and body swap video AI: Recast vs Sume's avatar face swap
Sume has two ways to put someone else in a video: H3 Max Recast swaps the person from a photo, Beta Face Swap applies a ready avatar's face. Which to use.
- Face swap vs H3 Max Recast: which swaps the person in a video?
Sume offers two ways to put a different person in a video: Avatar Face Swap (Beta) and H3 Max Recast. Inputs, length limits, audio and price side by side.
Written by Sume