Flux TTS pause markers: 8 per request, 500-3000 ms, and Sume scripts
Flux TTS allows up to 8 pause markers of 500 to 3000 ms per request, batch only. Plan a script around that limit and see what Sume gives you for pacing.

A Flux TTS request can carry at most 8 inline pause markers, each from 500 to 3000 ms in 100 ms steps, and only on batch (REST) requests. That is a script-planning limit: a long narration with a pause after every paragraph will not fit. Sume's TTS request takes up to 20000 characters of transcript, and its sentence segmentation returns timings you can use to place silence yourself.
Sources: the Deepgram changelog entry of September 30, 2026 and Sume's API reference, read 2026-10-01.
What are the Flux pause limits?
| Limit | Value |
|---|---|
| Marker example | \{pause:1s\} |
| Duration range | 500 to 3000 ms |
| Step | 100 ms |
| Markers per request | Up to 8 |
| Availability | Batch (REST) requests |
| Speed with a pause | Capped at 1.15 |
How do I plan a script around 8 pauses?
Count your breaks before you write. Spend the eight markers on the pauses that carry meaning, such as before a reveal or after a price, and let punctuation handle the rest. If a script needs more, split it into several requests so each carries its own eight. Remember the speed cap: a pause in the text means speed cannot go above 1.15.
What does Sume give me for pacing?
Sume's schema has timestamps.words: true, which returns word-level timings on the completed job, and segmentation for sentence segments that are gapless: each segment ends where the next begins. The schema says segmentation requires timestamps.words: true. With those timings you can insert silence in your own edit rather than inside the request.
Speed is a [0.6, 1.5] multiplier in generation_config, and an optional emotion guide is available. The cited schema does not describe an inline pause marker, so this post does not claim one.
When is each approach the better fit?
Use inline markers when you want the silence generated with the voice and your script has few breaks. Use timings and an edit when you have many, or when the pauses depend on the picture. For a wider comparison of marker styles see adding pauses in text to speech.
Sources
Related posts
More in Developers
- Deepgram self-hosted 261001 drops Whisper: what Sume STT callers do
Deepgram Self-Hosted release 261001 removes Whisper support and asks you to move to Nova-3 first. What it means for Sume STT, which has no model to migrate.
- Deno fetch stops retrying fresh-connection POSTs: Sume keys
Deno now retries fetch transport errors only on reused connections. That cuts duplicate POSTs, but a paid Sume submit still needs an Idempotency-Key.
- Remove video speckle noise by API: median filter radius
Sume's video filter allowlists median, which replaces each pixel with the middle value of its neighbours. Radius runs 1 to 127; start at 1 and compare.
- Extract audio from MP4 to WAV or MP3: the detach defaults
Sume audio detach returns WAV (pcm_s16le, sample-exact) by default or MP3 at 128 kbps. Pick WAV for speech-to-text and timeline use, MP3 when file size matters.
Written by Sume