Pause between narration lines: TTS pause markers or audio concat?
Deepgram Flux TTS allows pause markers of 500 to 3000 ms, max 8 per request. Sume's timeline audio concat joins lines with no gap. Where to put the pause.

Put the pause where the audio is made, not where it is joined. Deepgram's Flux TTS supports pause markers of 500 to 3000 ms, but Sume's timeline audio concat joins clips sample by sample with no silence added at the seams, so a gap between two separately generated lines has to already exist inside one of the clips.
What Deepgram's pause markers allow
The Deepgram changelog entry dated September 30 describes pause markers in Flux TTS with fixed limits, plus inline IPA pronunciation and a cap on speed when pauses are used.
| Control | Limit |
|---|---|
| Pause length | 500 to 3000 ms in 100 ms steps |
| Pauses per request | Maximum 8 |
| Speed with pauses | Capped at 1.15 |
| Pronunciation | Inline IPA |
What the Sume concat does at a seam
Sume's timeline audio job with operation: "concat" takes 1 to 20 ordered parts and returns one gapless file. The docs describe the join as sample-domain with no re-synthesis and no silence at the seams. The result includes segments[] with index, start and duration_seconds, which are the offsets you use to re-base video starts. The job costs $0.01 flat and the default output is sample-exact wav.
The consequence is simple: two lines concatenated will play back to back. If you want a beat between them, generate it in the line, for example with a vendor pause marker.
Choosing where to pause
- Short beats inside a sentence or between two sentences of one line: use the pause markers your TTS supports, within its per-request cap.
- Many lines in one scene: keep each line a separate clip so you can re-time a single line, then concat once and read
segments[]for the offsets. - More than 8 pauses in a read: split the text across requests so that each stays under the cap.
A concat request with trimmed edges
A part can carry source_in and duration, which lets you trim breath or tail from a line before the join. Trimming is the opposite of padding: it tightens the rhythm. Preview the result, then adjust.
curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: narration-concat-001" \
-d '{
"operation": "concat",
"parts": [
{ "url": "https://media.sume.com/artifacts/artf_demo/line1.wav" },
{ "url": "https://media.sume.com/artifacts/artf_demo/line2.wav", "source_in": 0.1, "duration": 1.8 }
]
}'Sources
Related posts
More in Developers
- Performance Max text limits: 30, 90, 90 and 25 characters, counted
Performance Max headlines run 30 characters, long headlines 90, descriptions 90, business name 25. A short script counts generated copy before upload.
- Push or poll for a finished render: listen, webhook or jobs_wait
MCP 2026-07-28 adds subscriptions/listen. For a render that takes minutes, compare a listen stream, a signed webhook and jobs_wait, with a Python verifier.
- Python 3.10 is end of life: a stdlib Sume webhook verifier
Python 3.10 has reached end of life. A standard-library verifier for Sume's signed webhooks that refuses an empty secret and accepts rotated signatures.
- Recraft V4.1 Flash: median 1.3 s, p95 1.8 s. Set timeouts from p95
Recraft quotes a median of about 1.3 seconds and a p95 of 1.8 seconds for V4.1 Flash. How to turn latency claims into timeouts and polling for image APIs.
Written by Sume