Eleven v4 inline tags like [whispers] vs Sume's emotion guide field

Eleven v4 steers delivery with inline tags in the text. Sume TTS has a separate emotion string, plus speed and volume. How to port a tagged script.

5 min readSume
All posts

Eleven v4 puts delivery direction inside the text, with tags such as [pause] and [whispers]. Sume TTS keeps the text clean and puts direction in a separate generation_config object with an emotion string, a speed multiplier and a volume multiplier. If you are porting a tagged Eleven script, strip the tags from the transcript and translate them into those three controls. Do not leave bracket tags in a Sume transcript; Sume documents no tag syntax for the text.

The two approaches differ in where the direction lives, and that decides how you move a script.

What does each side document?

On the Eleven v4 page, read 2026-10-04, ElevenLabs says you can direct delivery with inline tags, and names [pause], [long pause], [whispers] and [excited]. In the Sume OpenAPI spec, generation_config is an object with volume (0.5 to 2), speed (0.6 to 1.5) and emotion, a string of 1 to 64 characters described as an optional emotion guide.

Delivery controls, vendor pages read 2026-10-04
NeedEleven v4 (page)Sume TTS 1.0 (spec)
Whisper or excitementInline tags in the textemotion string, up to 64 characters
Pause[pause], [long pause]No tag syntax documented; use punctuation and sentence splits
PaceNot covered herespeed 0.6 to 1.5 for the whole request
LoudnessNot covered herevolume 0.5 to 2.0

How do you port a tagged script?

The key limit: Sume's controls apply to a request, not to a word. A [whispers] in the middle of a paragraph has no direct equivalent. Split the script at the point where the delivery changes and send each part as its own request with its own emotion.

  • Remove every bracket tag from the transcript.
  • Group lines by delivery and send one request per group.
  • Join the audio afterwards with a timeline audio concat.
  • Judge the result by ear; Sume does not promise a tag-for-tag match.

What does the request look like?

One group of lines with a single emotion guide:

import os, requests

H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}

r = requests.post("https://api.sume.com/v1/tts-1.0/generate",
    headers=H, timeout=60,
    json={
        "transcript": "It was quiet. Nobody moved.",
        "voice": {"id": os.environ["VOICE_ID"]},
        "language": "en",
        "generation_config": {"emotion": "hushed, close to the mic",
                              "speed": 0.9},
    })
r.raise_for_status()
print(r.json()["data"]["job"]["id"])

Where can you read more?

A similar port from another vendor is in Qwen Audio 3.0 inline tags vs Sume's emotion guide, and pacing to a time slot is in fitting a voiceover to a 30 second slot.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume