Grok TTS speech tags like laugh and whisper vs Sume's emotion field
xAI lets you write [laugh] and wrap text in whisper or singing tags. Sume TTS documents no tag grammar, only emotion, speed and volume in generation_config.

Can I put [laugh] or a whisper tag in a Sume TTS transcript?
Not as a documented control. The xAI text-to-speech docs define a tag grammar that lives inside the text; the Sume TTS 1.0 contract defines none, and steers delivery through three separate fields instead. Test any bracket syntax on a short sample before you rely on it, because Sume does not promise it will be interpreted.
This post lays the two side by side so you know which script to write for which API.
Which tags does xAI document?
The xAI page splits tags into two kinds. Inline tags drop a vocal event at a point in the text, and wrapping tags change how a stretch of text is spoken.
- Inline: [pause], [long-pause], [laugh], [cry], [sigh], [cough], [throat-clear], [gasp], [inhale], [exhale].
- Wrapping: <whisper>, <singing>, <fast>, <slow>, <soft>, <loud>.
- The page also states that all voices can speak every supported language.
What does Sume give you instead?
Sume TTS 1.0 has an optional generation_config object with three controls, all documented in the OpenAPI contract. They apply to the whole request, not to a span of words.
| Control | xAI TTS | Sume TTS 1.0 |
|---|---|---|
| Whisper or singing for a phrase | Wrapping tags | Not available per phrase |
| Laugh, sigh, cough at a point | Inline tags | No documented tag |
| Speed | <fast> and <slow> wrapping tags | generation_config.speed, 0.6 to 1.5 |
| Loudness | <soft> and <loud> wrapping tags | generation_config.volume, 0.5 to 2.0 |
| Overall mood | Not a separate field on the page | generation_config.emotion, free text up to 64 characters |
| Pronunciation | Not covered on the page | pronunciation_dict_id |
How do you get phrase-level control on Sume anyway?
Split the script into separate TTS jobs, one per stretch that needs a different delivery, and set generation_config on each. A whispered aside becomes its own request with a quiet volume and a soft emotion string; the main read uses defaults. Then join the results with timeline audio concat, which is gapless, so add any pause you want in the text itself with punctuation or in the edit.
Request sentence timings with timestamps.words and segmentation.mode set to sentence if you want per-sentence slices; the contract says those slices need a wav or raw container, and mp3 returns timings without slice files.
How do you port a tagged xAI script to plain text for Sume?
Strip the tags first, then re-create the delivery with per-job settings. The snippet below removes the xAI inline and wrapping tags listed above and prints the plain transcript you would send as the Sume transcript field.
import re
INLINE = r"\[(?:pause|long-pause|laugh|cry|sigh|cough|throat-clear|gasp|inhale|exhale)\]"
WRAP = r"</?(?:whisper|singing|fast|slow|soft|loud)>"
def to_plain(text):
text = re.sub(INLINE, "", text)
text = re.sub(WRAP, "", text)
return re.sub(r"\s{2,}", " ", text).strip()
print(to_plain("Hello [laugh] <whisper>this is a secret</whisper>."))Is the difference worth switching for?
If your content depends on laughs, sighs and singing inside one read, xAI's tags are the closer match and Sume's per-request knobs will feel blunt. If your content is narration, ads and explainer voiceover, three request-level controls plus a stable voice are usually enough, and the emotion, speed and volume guide covers practical values.
Whichever you choose, keep the tagged script separate from the plain transcript. Tags that one engine treats as stage directions can be read aloud as literal words by another.
Sources
Related posts
More in Comparisons
- H3 Max Recast vs Genjutsu: which person swap to call on Sume
Sume lists two person-swap video rows. Recast: 1-4 people in a 5-30 s clip at 768p or 1080p. Genjutsu: 1-8 images at 480p or 720p. How to choose.
- Hedra 402 INSUFFICIENT_BALANCE vs Sume insufficient_credits
Hedra returns 402 INSUFFICIENT_BALANCE until you add funds; Sume returns 402 insufficient_credits. How each wallet check behaves and how to preflight a job.
- Hedra Avatar needs start frame and audio; Sume takes a handle and scri
Hedra Avatar generates from a start frame plus an audio track, up to 10 minutes. Sume's talking-video takes an avatar handle and a script, up to 60 seconds.
- Hedra Character 3 allows 10 minutes; what is the Sume equivalent?
Hedra lists Character 3 at up to 10 minutes. Sume's Avatar 1.0 takes 4-60 seconds per job, so 10 minutes means ten or more jobs joined on a timeline.
Written by Sume