Gemini TTS short pause and breath tags vs Sume sentence segments

Gemini 3.8 Flash TTS takes inline tags like short pause and breath. Sume lists no inline tags; here is how to control pacing with segments and speed.

4 min readSume
All posts

Gemini 3.8 Flash TTS lets you place point-in-time events in the text with angle-bracket tags such as <short pause>, <long pause>, <breath>, <sigh> and <cough> (read 2026-10-07). Sume's TTS 1.0 contract documents no inline tag syntax. On Sume you shape pacing with sentence-level jobs, the speed and emotion controls in generation_config, and the sentence timings it returns.

Two ways to direct delivery on Gemini

The Gemini docs split control into two tiers. Turn-level delivery goes in speech_metadata.style for a sustained emotion or pace. Point-in-time events use inline tags. Capitalization adds emphasis, punctuation and ellipses create hesitation, and a pipe-wrapped reaction such as |reaction| handles backchannel in two-speaker dialogue.

Pacing controls on Gemini 3.8 Flash TTS (read 2026-10-07)
ControlWhere it goesEffect
speech_metadata.styleRequest fieldSustained emotion and pacing for the turn
<short pause>, <long pause>Inline in textA pause at that point
<breath>, <sigh>, <cough>, <laugh>Inline in textA vocal event at that point
Capitals, ellipsesInline in textEmphasis and hesitation

What Sume gives you instead

The Sume TTS 1.0 schema has generation_config with speed (0.6 to 1.5), volume (0.5 to 2) and a free-text emotion guide of up to 64 characters. There is a pronunciation_dict_id for words that read wrong. The docs list nothing like <short pause>, so do not assume a tag is honored; test a short line first and listen.

For a longer gap between lines, split the script into one job per sentence or paragraph and join the clips on a timeline. Requesting timestamps.words: true with segmentation.mode: sentence returns gapless segments, and the cut sits 70 ms after each sentence's last word unless you change boundary_lead_ms (0 to 500). The next segment absorbs the natural pause.

  • A beat before a price: end the sentence before it, and let the cut keep the pause.
  • A slower read for a disclaimer: lower speed on that job only.
  • A breath or laugh sound: not documented on Sume; if it matters, record it elsewhere or pick the Gemini tags.

Request with segments

Segment audio slices need a wav or raw container. With mp3 you still get the timings, but no per-segment audio_url.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: pause-test-001" \
  -d '{
    "transcript": "Our price is simple. Twelve dollars a month.",
    "avatar_handle": "speaker",
    "language": "en",
    "generation_config": {"speed": 0.9},
    "output_format": {"container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100},
    "timestamps": {"words": true},
    "segmentation": {"mode": "sentence", "boundary_lead_ms": 150}
  }'

Which to choose

If you need mid-sentence vocal events, Gemini's tags are a feature Sume does not document. If you need predictable sentence-level cuts that line up with captions and scene changes, Sume's segments are built for that. Sume does not call Gemini, so you pick one or run the two as separate vendors.

Sources

Related posts

More in Models

All Models posts

Written by Sume