Eleven v4 inline tags like [whispers] vs Sume's emotion guide field
Eleven v4 steers delivery with inline tags in the text. Sume TTS has a separate emotion string, plus speed and volume. How to port a tagged script.

Eleven v4 puts delivery direction inside the text, with tags such as [pause] and [whispers]. Sume TTS keeps the text clean and puts direction in a separate generation_config object with an emotion string, a speed multiplier and a volume multiplier. If you are porting a tagged Eleven script, strip the tags from the transcript and translate them into those three controls. Do not leave bracket tags in a Sume transcript; Sume documents no tag syntax for the text.
The two approaches differ in where the direction lives, and that decides how you move a script.
What does each side document?
On the Eleven v4 page, read 2026-10-04, ElevenLabs says you can direct delivery with inline tags, and names [pause], [long pause], [whispers] and [excited]. In the Sume OpenAPI spec, generation_config is an object with volume (0.5 to 2), speed (0.6 to 1.5) and emotion, a string of 1 to 64 characters described as an optional emotion guide.
| Need | Eleven v4 (page) | Sume TTS 1.0 (spec) |
|---|---|---|
| Whisper or excitement | Inline tags in the text | emotion string, up to 64 characters |
| Pause | [pause], [long pause] | No tag syntax documented; use punctuation and sentence splits |
| Pace | Not covered here | speed 0.6 to 1.5 for the whole request |
| Loudness | Not covered here | volume 0.5 to 2.0 |
How do you port a tagged script?
The key limit: Sume's controls apply to a request, not to a word. A [whispers] in the middle of a paragraph has no direct equivalent. Split the script at the point where the delivery changes and send each part as its own request with its own emotion.
- Remove every bracket tag from the transcript.
- Group lines by delivery and send one request per group.
- Join the audio afterwards with a timeline audio concat.
- Judge the result by ear; Sume does not promise a tag-for-tag match.
What does the request look like?
One group of lines with a single emotion guide:
import os, requests
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
r = requests.post("https://api.sume.com/v1/tts-1.0/generate",
headers=H, timeout=60,
json={
"transcript": "It was quiet. Nobody moved.",
"voice": {"id": os.environ["VOICE_ID"]},
"language": "en",
"generation_config": {"emotion": "hushed, close to the mic",
"speed": 0.9},
})
r.raise_for_status()
print(r.json()["data"]["job"]["id"])
Where can you read more?
A similar port from another vendor is in Qwen Audio 3.0 inline tags vs Sume's emotion guide, and pacing to a time slot is in fitting a voiceover to a 30 second slot.
Sources
Related posts
More in Comparisons
- Eleven v4 Turbo 150 ms first speech vs Sume async TTS jobs
ElevenLabs quotes about 150 ms to first speech for v4 Turbo. Sume TTS is an async job, built for finished narration. Which one fits your use?
- ElevenLabs Music for a TV spot: two pages that do not agree
ElevenLabs music page says self-serve plans exclude film, TV and games; the API docs say cleared for film and TV. Get this in writing before a broadcast ad.
- ElevenLabs Music terms: restricted industries and banned inputs
Before an ad brief goes into ElevenLabs Music: the terms list restricted industries and prompt inputs you cannot send. What Sume does and does not enforce.
- ElevenLabs TTS on fal at $0.10 per 1,000 characters vs direct and Sume
fal lists ElevenLabs TTS at $0.10 per 1,000 characters, ElevenLabs lists $0.08 for v3, Sume is $0.0475. A 2,000-character script: $0.20, $0.16, $0.095.
Written by Sume