Gemini TTS short pause and breath tags vs Sume sentence segments
Gemini 3.8 Flash TTS takes inline tags like short pause and breath. Sume lists no inline tags; here is how to control pacing with segments and speed.

Gemini 3.8 Flash TTS lets you place point-in-time events in the text with angle-bracket tags such as <short pause>, <long pause>, <breath>, <sigh> and <cough> (read 2026-10-07). Sume's TTS 1.0 contract documents no inline tag syntax. On Sume you shape pacing with sentence-level jobs, the speed and emotion controls in generation_config, and the sentence timings it returns.
Two ways to direct delivery on Gemini
The Gemini docs split control into two tiers. Turn-level delivery goes in speech_metadata.style for a sustained emotion or pace. Point-in-time events use inline tags. Capitalization adds emphasis, punctuation and ellipses create hesitation, and a pipe-wrapped reaction such as |reaction| handles backchannel in two-speaker dialogue.
| Control | Where it goes | Effect |
|---|---|---|
| speech_metadata.style | Request field | Sustained emotion and pacing for the turn |
| <short pause>, <long pause> | Inline in text | A pause at that point |
| <breath>, <sigh>, <cough>, <laugh> | Inline in text | A vocal event at that point |
| Capitals, ellipses | Inline in text | Emphasis and hesitation |
What Sume gives you instead
The Sume TTS 1.0 schema has generation_config with speed (0.6 to 1.5), volume (0.5 to 2) and a free-text emotion guide of up to 64 characters. There is a pronunciation_dict_id for words that read wrong. The docs list nothing like <short pause>, so do not assume a tag is honored; test a short line first and listen.
For a longer gap between lines, split the script into one job per sentence or paragraph and join the clips on a timeline. Requesting timestamps.words: true with segmentation.mode: sentence returns gapless segments, and the cut sits 70 ms after each sentence's last word unless you change boundary_lead_ms (0 to 500). The next segment absorbs the natural pause.
- A beat before a price: end the sentence before it, and let the cut keep the pause.
- A slower read for a disclaimer: lower
speedon that job only. - A breath or laugh sound: not documented on Sume; if it matters, record it elsewhere or pick the Gemini tags.
Request with segments
Segment audio slices need a wav or raw container. With mp3 you still get the timings, but no per-segment audio_url.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: pause-test-001" \
-d '{
"transcript": "Our price is simple. Twelve dollars a month.",
"avatar_handle": "speaker",
"language": "en",
"generation_config": {"speed": 0.9},
"output_format": {"container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100},
"timestamps": {"words": true},
"segmentation": {"mode": "sentence", "boundary_lead_ms": 150}
}'Which to choose
If you need mid-sentence vocal events, Gemini's tags are a feature Sume does not document. If you need predictable sentence-level cuts that line up with captions and scene changes, Sume's segments are built for that. Sume does not call Gemini, so you pick one or run the two as separate vendors.
Sources
Related posts
More in Models
- Gemini Omni Flash 1.1 API cost per clip, 360p to 4K
Gemini Omni Flash 1.1 on Sume is $0.0375 to $0.375 per second. An 8-second 720p clip is $1.00; 100 clips are $100. Full table by resolution.
- gemini-omni-flash-preview: Oct 22 shutdown on the table, not Sep 30
Google's release note says gemini-omni-flash-preview is deprecated on Sept 30, 2026, but its deprecations table now lists Oct 22. What to pin and what to test.
- Three Gemini TTS preview ids shut down Nov 17: what replaces them
gemini-3.1-flash-tts-preview and both 2.5 TTS previews are listed for shutdown on Nov 17, 2026. Google names two replacements; here is the price per 10 seconds.
- Gemini voice remixing is 'coming soon': what Sume TTS can change today
Google says Gemini voice remixing for timbre, pitch, pace and accent is coming soon. Sume TTS today has speed, volume and an emotion string only.
Written by Sume