Grok TTS speech tags like laugh and whisper vs Sume's emotion field

xAI lets you write [laugh] and wrap text in whisper or singing tags. Sume TTS documents no tag grammar, only emotion, speed and volume in generation_config.

5 min readSume
All posts

Can I put [laugh] or a whisper tag in a Sume TTS transcript?

Not as a documented control. The xAI text-to-speech docs define a tag grammar that lives inside the text; the Sume TTS 1.0 contract defines none, and steers delivery through three separate fields instead. Test any bracket syntax on a short sample before you rely on it, because Sume does not promise it will be interpreted.

This post lays the two side by side so you know which script to write for which API.

Which tags does xAI document?

The xAI page splits tags into two kinds. Inline tags drop a vocal event at a point in the text, and wrapping tags change how a stretch of text is spoken.

  • Inline: [pause], [long-pause], [laugh], [cry], [sigh], [cough], [throat-clear], [gasp], [inhale], [exhale].
  • Wrapping: <whisper>, <singing>, <fast>, <slow>, <soft>, <loud>.
  • The page also states that all voices can speak every supported language.

What does Sume give you instead?

Sume TTS 1.0 has an optional generation_config object with three controls, all documented in the OpenAPI contract. They apply to the whole request, not to a span of words.

Delivery controls, xAI docs and Sume OpenAPI (read 2026-10-02)
ControlxAI TTSSume TTS 1.0
Whisper or singing for a phraseWrapping tagsNot available per phrase
Laugh, sigh, cough at a pointInline tagsNo documented tag
Speed<fast> and <slow> wrapping tagsgeneration_config.speed, 0.6 to 1.5
Loudness<soft> and <loud> wrapping tagsgeneration_config.volume, 0.5 to 2.0
Overall moodNot a separate field on the pagegeneration_config.emotion, free text up to 64 characters
PronunciationNot covered on the pagepronunciation_dict_id

How do you get phrase-level control on Sume anyway?

Split the script into separate TTS jobs, one per stretch that needs a different delivery, and set generation_config on each. A whispered aside becomes its own request with a quiet volume and a soft emotion string; the main read uses defaults. Then join the results with timeline audio concat, which is gapless, so add any pause you want in the text itself with punctuation or in the edit.

Request sentence timings with timestamps.words and segmentation.mode set to sentence if you want per-sentence slices; the contract says those slices need a wav or raw container, and mp3 returns timings without slice files.

How do you port a tagged xAI script to plain text for Sume?

Strip the tags first, then re-create the delivery with per-job settings. The snippet below removes the xAI inline and wrapping tags listed above and prints the plain transcript you would send as the Sume transcript field.

import re

INLINE = r"\[(?:pause|long-pause|laugh|cry|sigh|cough|throat-clear|gasp|inhale|exhale)\]"
WRAP = r"</?(?:whisper|singing|fast|slow|soft|loud)>"

def to_plain(text):
    text = re.sub(INLINE, "", text)
    text = re.sub(WRAP, "", text)
    return re.sub(r"\s{2,}", " ", text).strip()

print(to_plain("Hello [laugh] <whisper>this is a secret</whisper>."))

Is the difference worth switching for?

If your content depends on laughs, sighs and singing inside one read, xAI's tags are the closer match and Sume's per-request knobs will feel blunt. If your content is narration, ads and explainer voiceover, three request-level controls plus a stable voice are usually enough, and the emotion, speed and volume guide covers practical values.

Whichever you choose, keep the tagged script separate from the plain transcript. Tags that one engine treats as stage directions can be read aloud as literal words by another.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume