Eleven v4 drops the native accent; Sume keeps a language tag per voice

ElevenLabs' docs say v4 gives fluent target-language speech, not a preserved accent. Sume tags each voice with a language and checks for mismatch.

5 min readSume
All posts

ElevenLabs' Eleven v4 docs say v4 produces fluent speech in the target language rather than preserving the speaker's native accent (read 2026-10-03). So a cloned voice speaking another language will sound like a native speaker of it, not like the original speaker with an accent. Sume works differently on the surface: each voice carries one language tag, you send a language code with the transcript, and a mismatch between the two returns a warning before any charge.

Side by side

The ElevenLabs statements are from its model docs, and the Sume ones from the MCP tools and gates page.

Behaviour as described on each vendor's pages, read 2026-10-03
QuestionEleven v4 docsSume tts_create
Cross-language outputFluent target-language speech, no preserved native accentVoice has one language tag; send matching language
Language omittedNot stated on the pageProvider assumes English
Voice and language disagreeNot stated on the pageWarning, no job and no charge until confirmed

Which one you want

If a spokesperson should keep their accent across languages, neither page promises that. The ElevenLabs page says the opposite. For a multilingual channel, the safer route on Sume is one voice per language, each tagged, so the voice and the transcript agree. Listen to a sample before a batch in either service.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume