Text to speech add pause: Flux, HeyGen and Gemini markers vs Sume
Deepgram Flux, HeyGen and Gemini TTS each use their own pause marker. Sume TTS has none, so use punctuation, speed, or a declared silence in Timeline.

To add a pause in text to speech you normally put a marker in the text, and each vendor's marker is different. Sume TTS documents no pause marker in its request, so on Sume you shape pauses with punctuation, the speed control, or by generating separate lines and joining them.
Vendor syntax is from the Deepgram, HeyGen and Gemini pages, read 2026-09-30; Sume facts are from the OpenAPI document and Timeline docs.
What pause markers do other vendors use?
| Vendor | Marker | Limits stated |
|---|---|---|
| Deepgram Flux TTS | \{pause:1s\} | 500 to 3000 ms in 100 ms steps, up to 8 per request, batch REST |
| HeyGen | <break time="0.5s"/> | Max 5 seconds per pause, professional voice clones |
| Gemini TTS | <short pause> inside the transcript | Point-in-time inline tag |
| Sume TTS 1.0 | None documented | Speed 0.6 to 1.5, pronunciation dictionary id |
What controls does Sume TTS have?
The request has an optional generation_config with volume, speed and emotion controls, and an optional pronunciation_dict_id. Speed is a multiplier in [0.6, 1.5]. None of these inserts a measured silence, so they change delivery, not gap length. Emotion and speed are covered in TTS with emotion, speed and volume.
How do I get a longer gap on Sume?
Generate each sentence or paragraph as its own TTS job and join the files; Timeline audio joins Sume-hosted audio into one gapless file, which means the join itself adds no gap. For a pause of a known length you can place clips on a Timeline spine with later start times. Timeline's audio.mode value silence declares an output length with no audio file, which is for silent videos rather than a silence clip to splice. See join voiceover parts in one render.
Will punctuation alone work?
Commas, periods and paragraph breaks affect phrasing in most voices, but the Sume docs do not promise a duration for any of them. Listen to the result and adjust, and do not paste another vendor's marker into a Sume request: the text would be read as words.
Sources
Related posts
More in Use cases
- TikTok 3-minute videos: cut a longer clip to 180 seconds via API
TikTok's page says all creators can post 3-minute videos, some 5 or 10. Sume video-trim writes a 180 s MP4 from a longer clip so any creator can post it.
- TikTok AI voiceover label: generic TTS vs a real person's voice
TikTok's 2026-H2 guidelines: generic text-to-speech narration needs no label, but AI audio that mimics a real person's voice does. How that maps to Sume voices.
- TikTok AI label policy: does anime or cartoon video need one?
TikTok counts anime and cartoons as AI content but says artistic styles need no disclosure. Realistic-looking people or scenes still do. The edge case.
- TikTok API 10-minute video limit: trim a longer video with Sume
TikTok's media guide says a developer can send at most 10 minutes through initialize upload. Sume video-trim cuts up to 900 s from a 1800 s source into an MP4.
Written by Sume