Eleven v4 stacked tags: direct emotion without them
ElevenLabs v4 adds stackable expression tags and 10-second cloning. Sume's docs list no clone route; here is how to direct a voice via tts_create.

TechCrunch reported on September 28, 2026 that ElevenLabs v4 can stack inline expression tags and clone a voice from 10 seconds of audio. Sume's docs list a tts_create tool but no voice-cloning route, so on Sume you direct delivery through the script itself and check the tool's contract first.
What v4 changed, per the report
- Clones a voice from 10 seconds of audio.
- Languages up from 70 to 90.
- Improved Japanese, Brazilian Portuguese, Mandarin and Cantonese.
- Inline tags can be stacked, and the model follows the sequence.
What Sume documents
The hosted MCP lists tts_create among paid generation tools, and tts_source_get and tts_source_verify_spine as free reads for accepted-script checks. The docs I read do not describe tag syntax, cloning or a voice list, so do not assume a v4-style tag works. Call tools_schema with name: "tts_create" to read the live contract before sending a script.
Call tools_schema with name "tts_create" and list every field
that controls voice, language and script before submitting.Direct a voice without tags
Delivery is mostly decided by the text. These habits work with any speech model and need no special syntax:
- Write short sentences; one idea each.
- Use punctuation for pacing: commas for breath, full stops for finality.
- Spell out numbers and abbreviations the way you want them said.
- Render a line at a time, then keep the takes you like; a wave of parallel creates is what
script_runand batchjobs_waitare for. - Listen, then rewrite the script instead of fighting the output.
Talking faces
If the voice drives an on-camera person, the Sume docs say a talking face is a Fabric clip made from an accepted still plus the TTS audio. Video models do not lip-sync to narration added later.
Sources
Related posts
More in Models
- eleven_v4 vs eleven_v4_turbo: model IDs, endpoints, which to pick
ElevenLabs lists eleven_v4 for expressive speech with cloning in 90+ languages and eleven_v4_turbo at about 100 ms median latency. Which fits a video pipeline.
- ElevenLabs languages: Flash v2.5 has 32, Multilingual v2 29, v4 90+
ElevenLabs lists 32 languages for Flash v2.5, 29 for Multilingual v2 and 90+ for v4 and v4 Turbo. Check your markets against the model, then log it per job.
- Gemini 3.8 Flash TTS tops Hume's VoiceEQ board: run your own test
Hume's blog lists Gemini 3.8 Flash TTS atop its Real-World VoiceEQ board. Why a vendor-run board is only a lead, and how to run a blind A/B on your script.
- Gemini Omni Flash GA: extension and interpolation vs Sume's inputs
Gemini Omni Flash is generally available with extension and interpolation between images. What Sume's gemini-omni-flash-1.1 catalog entry documents instead.
Written by Sume