Directing delivery: Eleven v4 audio tags vs Gemini TTS style
Eleven v4 puts direction inline as audio tags like [laughs]; Gemini TTS adds a separate style field. Sume sends transcript, voice and language only.

They solve the same problem in two places. Eleven v4 keeps direction inside the script text as inline audio tags, so one string carries words and delivery. Gemini TTS has a style field for overall delivery and also accepts inline angle-bracket tags, so you can set a baseline and then adjust a line. Sume's TTS Router has neither: the documented request takes a transcript, a voice, a language and output options, and the delivery comes from the voice and the words.
Which is easier depends on your pipeline. Inline tags travel with the text, which suits script-driven tools. A separate field is cleaner when delivery is a setting you vary across many scripts.
What the vendors state
| Item | Eleven v4 | Gemini TTS |
|---|---|---|
| Delivery control | Inline audio tags | style field plus inline angle-bracket tags |
| Examples given | [laughs], [said angrily in French accent], [light rain], [phone buzzing] | Examples not reproduced here; see the Google page |
| Pronunciation | Better IPA phoneme support | Not stated on the page read |
| Speakers | Multi-speaker dialogue | Up to 2 speakers per request with prebuilt voices |
| Output formats | MP3, WAV/PCM, mu-law (product page) | WAV default, raw PCM streaming, mu-law/A-law at 8000-24000 Hz |
Notice what a tag can and cannot do
- Eleven's examples mix performance cues ([laughs]) with scene sound ([light rain], [phone buzzing]). That means tags can change the audio bed, not only the voice. Check your usage rules before you ship a track that contains generated ambience.
- A tag is a request, not a guarantee. Neither vendor page read for this post promises that every tag renders the same on every voice, so render and listen.
- Keep tags out of the text you also use for captions. Strip them or you will burn the brackets onto the video.
Where Sume stands
Sume does not ship Eleven v4 or Gemini TTS. The TTS Router serves Cartesia Sonic (sonic-3.6, sonic-3.5, sonic-3, sonic-latest, sonic-preview). The request fields the reference lists are model, transcript or transcript_source, a voice selector, language, output_format, timestamps, segmentation and webhook fields. No style or tag field is documented, and this post makes no claim about how Sonic handles bracket tags.
A practical consequence: if you direct delivery with tags today, keep the tag layer in a separate column of your script file so the same script can feed a tag-aware engine or a plain one.
Sources
Related posts
More in Comparisons
- Eleven v4 IPA support vs Sume's pronunciation_dict_id for brand names
ElevenLabs says Eleven v4 improves IPA phoneme support. Sume TTS has an optional pronunciation_dict_id. How each helps with brand names today.
- Voice agent TTS latency: Eleven v4 Turbo 150ms vs Cartesia Sonic
Eleven v4 Turbo claims about 150ms to first speech; Cartesia Sonic-3.6 claims under 90ms. Not like-for-like; Sume TTS is not for live agents.
- Eleven v4 or Sonic 3.6 for ad voiceover through an API
Eleven v4 lists $0.022 per 1K characters, Sonic 3.6 $0.038 at list. Sume serves Sonic 3.6 only, at list x1.25. Here is how to pick for ad reads.
- ElevenLabs Ads Engine localizes ads in 50+ languages; what Sume covers
Ads Engine pulls ads from Google, Meta and LinkedIn and localizes them. Sume has no ad-account link, but burns translated captions at $0.20 a language.
Written by Sume