Eleven v4 audio tags [laughs]: how to use them, and Sume's controls
Eleven v4 reads inline tags such as [laughs] and [light rain] in the text. Sume's TTS schema has no tag syntax; it has speed, volume and emotion controls.

To use Eleven v4 audio tags, write them inline in the text, in square brackets, where the sound or delivery should happen. ElevenLabs' launch post gives examples: [laughs], [said angrily in French accent], [light rain] and [phone buzzing]. Sume's text-to-speech routes do not list Eleven models and document no tag syntax; they take plain transcript text with speed, volume and emotion controls.
The Eleven facts are from ElevenLabs' launch post, read 2026-09-29. The Sume facts are from the TTS schemas in the Sume API reference.
What tags does ElevenLabs show?
ElevenLabs also says you can direct delivery in natural language. These are its examples, and we did not test them.
| Kind of cue | Example tag |
|---|---|
| Vocal reaction | [laughs] |
| Delivery and accent | [said angrily in French accent] |
| Ambient sound | [light rain] |
| Object sound | [phone buzzing] |
Does Sume support inline tags?
Sume documents none. POST /v1/tts-router/generate requires a model that is one of sonic-3.6, sonic-3.5, sonic-3, sonic-latest or sonic-preview, and TTS 1.0 has no engine picker. Neither schema describes bracket tags, so do not assume [laughs] is a control. Test any tag on your own text before it goes into a script.
What controls does Sume give for delivery?
| Field | Range |
|---|---|
generation_config.speed | Multiplier from 0.6 to 1.5 |
generation_config.volume | Multiplier from 0.5 to 2.0 |
generation_config.emotion | A string of up to 64 characters |
pronunciation_dict_id | Optional pronunciation dictionary id |
curl -X POST https://api.sume.com/v1/tts-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: delivery-001" \
-d '{
"model": "sonic-3.6",
"transcript": "That is the news I have heard all week.",
"voice": { "id": "voi_..." },
"language": "en",
"generation_config": { "speed": 0.95, "emotion": "delighted" }
}'How do I add real sounds around speech?
A tag that makes rain or a phone buzz is inside the model. With Sume, speech is one file and sounds are another: generate the speech, then mix a bed under it with the Timeline soundtrack field. See text to speech pronunciation for the one text-level control Sume does document.
How do I write a script when tags are not available?
Put the reaction in the words and the direction in the fields. Write a line the speaker would say with a laugh in it, set generation_config.emotion to a short guide such as “amused”, and nudge speed for pacing. For a line that needs a different feel, send it as its own job with its own emotion.
ElevenLabs also says Eleven v4 covers more than 90 languages. Which languages a Sume voice speaks is set by the voice and the language field, so check the voice you plan to use.
What should I test before relying on a tag?
Run one short line with the tag and one without, on the same voice, and listen to both. If the tag is spoken aloud or ignored, remove it. Keep the result of that test with your script so the next editor does not reintroduce it.
Sources
Related posts
More in Models
- FLUX.2 dev commercial use: what the model card says
FLUX.2 [dev] weights carry the FLUX non-commercial license, yet the card says outputs can be used commercially as that license describes. The exact wording.
- FLUX.2 flex text rendering: prompt tips and the Sume model id
Black Forest Labs positions FLUX.2 [flex] for text and fine detail. How its typography guidance reads, and how to call black-forest-labs/flux.2-flex on Sume.
- FLUX.2 hex color prompt: how to write it, with an API call
Black Forest Labs says FLUX.2 matches hex codes in prompts. The syntax it documents, the limits it admits, and a flux.2-pro call on Sume.
- FLUX.2 pro vs Nano Banana Pro API: request specs side by side
Request parameters that differ between black-forest-labs/flux.2-pro and google/nano-banana-pro on Sume: resolution tiers, aspect ratios, references and price.
Written by Sume