Directing delivery: Eleven v4 audio tags vs Gemini TTS style

Eleven v4 puts direction inline as audio tags like [laughs]; Gemini TTS adds a separate style field. Sume sends transcript, voice and language only.

4 min readSume
All posts

They solve the same problem in two places. Eleven v4 keeps direction inside the script text as inline audio tags, so one string carries words and delivery. Gemini TTS has a style field for overall delivery and also accepts inline angle-bracket tags, so you can set a baseline and then adjust a line. Sume's TTS Router has neither: the documented request takes a transcript, a voice, a language and output options, and the delivery comes from the voice and the words.

Which is easier depends on your pipeline. Inline tags travel with the text, which suits script-driven tools. A separate field is cleaner when delivery is a setting you vary across many scripts.

What the vendors state

Vendor pages, read 2026-10-05
ItemEleven v4Gemini TTS
Delivery controlInline audio tagsstyle field plus inline angle-bracket tags
Examples given[laughs], [said angrily in French accent], [light rain], [phone buzzing]Examples not reproduced here; see the Google page
PronunciationBetter IPA phoneme supportNot stated on the page read
SpeakersMulti-speaker dialogueUp to 2 speakers per request with prebuilt voices
Output formatsMP3, WAV/PCM, mu-law (product page)WAV default, raw PCM streaming, mu-law/A-law at 8000-24000 Hz

Notice what a tag can and cannot do

  • Eleven's examples mix performance cues ([laughs]) with scene sound ([light rain], [phone buzzing]). That means tags can change the audio bed, not only the voice. Check your usage rules before you ship a track that contains generated ambience.
  • A tag is a request, not a guarantee. Neither vendor page read for this post promises that every tag renders the same on every voice, so render and listen.
  • Keep tags out of the text you also use for captions. Strip them or you will burn the brackets onto the video.

Where Sume stands

Sume does not ship Eleven v4 or Gemini TTS. The TTS Router serves Cartesia Sonic (sonic-3.6, sonic-3.5, sonic-3, sonic-latest, sonic-preview). The request fields the reference lists are model, transcript or transcript_source, a voice selector, language, output_format, timestamps, segmentation and webhook fields. No style or tag field is documented, and this post makes no claim about how Sonic handles bracket tags.

A practical consequence: if you direct delivery with tags today, keep the tag layer in a separate column of your script file so the same script can feed a tag-aware engine or a plain one.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume