Laugh and sigh tags in TTS: Gemini 3.8 tags vs Sume's emotion field
Gemini 3.8 TTS takes inline tags such as laugh, sigh and breath. Sume TTS documents an emotion hint, speed and volume instead; test tags on one short job.

Gemini 3.8 TTS documents inline vocal tags in the transcript, and its launch post describes scripted vocal bursts and backchanneling through non-verbal cues. Sume TTS documents three delivery controls instead: speed, volume and a short emotion string. Nothing in the Sume OpenAPI promises inline tags, so test one short job before relying on them.
What does each side document?
The Gemini row is from Google's speech generation page, read 2026-10-09. The Sume row is from the OpenAPI generation_config schema.
| Control | Gemini 3.8 TTS | Sume TTS 1.0 |
|---|---|---|
| Inline tags in the text | Documented examples: <cough>, <sigh>, <short pause>, <laugh>, <breath> | Not documented |
| Turn-level style | speech_metadata style, for example "whispered urgently" | emotion string, 1 to 64 characters |
| Pace | Not a numeric field on the page | speed 0.6 to 1.5 |
| Level | Not a numeric field on the page | volume 0.5 to 2.0 |
How do I test a tag on Sume without wasting a script?
Run a 40-character test with the tag in place and listen. A tag the engine does not understand may be read aloud as words, or dropped. That costs 1 cent, and the character count includes the tag, because the API bills spaces and punctuation as characters.
If the tag does not work, move the intent to the emotion field and the punctuation. A line such as "Oh. Well, that is that." with a short emotion guide carries a sigh without any markup.
What about stressing one word?
Gemini's own page shows capital letters as the way to put natural vocal stress on key words. That is a Gemini behaviour; the Sume docs do not describe it, so check it the same way, with one short job.
For a direct comparison of Gemini's turn style and Eleven's tags, see the style field comparison.
Sources
Related posts
More in Comparisons
- Generate at 1K and upscale, or generate 4K? Sume price check
Nano Banana 2.1 at 1K plus a Sume image upscale is $0.30, against $0.20 for native 4K. GPT Image 2.5 low plus upscale is $0.22 against $0.22 for high 4K.
- GPT Image 2.5 on Sume: low 1K $0.02475 vs high 4K $0.2225
GPT Image 2.5 spans 9x in price on Sume: $0.02475 for low at 1K, $0.055625 medium at 2K, $0.2225 high at 4K. Omit quality and the default is high.
- Griffin-Lite's 0.43 s latency: what a recorded avatar clip gives up
Tavus reports Griffin-Lite video latency of 0.43 s on average (0.27-0.59 s) in a research preview. A recorded clip on Sume is a job, so choose by use case.
- HappyHorse 1.0 at $0.15 a second: four Sume rows cost less at 720p
Runway Dev lists HappyHorse 1.0 at 15 credits per second at 720p and 30 at 1080p. Sume does not list it; four Sume rows are cheaper at 720p, three at 1080p.
Written by Sume