Eleven v4 audio tags [laughs]: how to use them, and Sume's controls

Eleven v4 reads inline tags such as [laughs] and [light rain] in the text. Sume's TTS schema has no tag syntax; it has speed, volume and emotion controls.

4 min readSume
All posts

To use Eleven v4 audio tags, write them inline in the text, in square brackets, where the sound or delivery should happen. ElevenLabs' launch post gives examples: [laughs], [said angrily in French accent], [light rain] and [phone buzzing]. Sume's text-to-speech routes do not list Eleven models and document no tag syntax; they take plain transcript text with speed, volume and emotion controls.

The Eleven facts are from ElevenLabs' launch post, read 2026-09-29. The Sume facts are from the TTS schemas in the Sume API reference.

What tags does ElevenLabs show?

ElevenLabs also says you can direct delivery in natural language. These are its examples, and we did not test them.

Examples from ElevenLabs' Eleven v4 post, read 2026-09-29.
Kind of cueExample tag
Vocal reaction[laughs]
Delivery and accent[said angrily in French accent]
Ambient sound[light rain]
Object sound[phone buzzing]

Does Sume support inline tags?

Sume documents none. POST /v1/tts-router/generate requires a model that is one of sonic-3.6, sonic-3.5, sonic-3, sonic-latest or sonic-preview, and TTS 1.0 has no engine picker. Neither schema describes bracket tags, so do not assume [laughs] is a control. Test any tag on your own text before it goes into a script.

What controls does Sume give for delivery?

From the Sume TTS request schema, read 2026-09-29.
FieldRange
generation_config.speedMultiplier from 0.6 to 1.5
generation_config.volumeMultiplier from 0.5 to 2.0
generation_config.emotionA string of up to 64 characters
pronunciation_dict_idOptional pronunciation dictionary id
curl -X POST https://api.sume.com/v1/tts-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: delivery-001" \
  -d '{
    "model": "sonic-3.6",
    "transcript": "That is the news I have heard all week.",
    "voice": { "id": "voi_..." },
    "language": "en",
    "generation_config": { "speed": 0.95, "emotion": "delighted" }
  }'

How do I add real sounds around speech?

A tag that makes rain or a phone buzz is inside the model. With Sume, speech is one file and sounds are another: generate the speech, then mix a bed under it with the Timeline soundtrack field. See text to speech pronunciation for the one text-level control Sume does document.

How do I write a script when tags are not available?

Put the reaction in the words and the direction in the fields. Write a line the speaker would say with a laugh in it, set generation_config.emotion to a short guide such as “amused”, and nudge speed for pacing. For a line that needs a different feel, send it as its own job with its own emotion.

ElevenLabs also says Eleven v4 covers more than 90 languages. Which languages a Sume voice speaks is set by the voice and the language field, so check the voice you plan to use.

What should I test before relying on a tag?

Run one short line with the tag and one without, on the same voice, and listen to both. If the tag is spoken aloud or ignored, remove it. Keep the result of that test with your script so the next editor does not reintroduce it.

Sources

Related posts

More in Models

All Models posts

Written by Sume