MiniMax <#1.5#> pause markup vs Sume TTS: no pause marker
MiniMax inserts pauses with <#x#> markers from 0.01 to 99.99 seconds. Sume TTS documents no pause marker; here is how to split a script and what to use instead.

Can I add a timed pause in Sume TTS the way MiniMax does?
Not with a documented marker. MiniMax lets you write a pause inside the text as <#x#>, where x is the duration in seconds from 0.01 to 99.99 with up to two decimals. The Sume TTS 1.0 request has no pause field and documents no marker syntax, so a script containing <#1.5#> is not a supported way to get a 1.5-second gap.
You can still shape pacing on Sume. The sections below split the work into what the text can do, what segmentation gives you and where an edit step is needed.
What does the Sume contract give you for pacing?
Three things. Punctuation and paragraph breaks in the transcript, which the voice reads as natural pauses. generation_config.speed from 0.6 to 1.5 to slow the whole read. And segmentation: with timestamps.words and segmentation.mode set to sentence, you receive gapless sentence segments, where boundary_lead_ms (0 to 500, default 70) sets how long after the last word each cut falls and the next segment absorbs the pause.
That last control moves a cut point; it does not synthesize silence. Treat it as a way to keep natural breath gaps inside the slices, not as a pause generator.
| Need | MiniMax T2A | Sume TTS 1.0 |
|---|---|---|
| Timed pause in the text | <#x#>, 0.01 to 99.99 s | No documented marker |
| Slower read overall | speed 0.5 to 2 | generation_config.speed 0.6 to 1.5 |
| Where sentences cut | Not covered here | segmentation.boundary_lead_ms 0 to 500, default 70 |
| Silent beat in video | Not covered here | Avatar video voice.type silence, duration required |
How do you port a script that already uses <#x#> markers?
Split at each marker, generate each stretch as its own job and keep the pause lengths in a list. The helper below returns pairs of text and the pause that followed it, so your edit step knows where gaps belong.
import re
def split_pauses(script):
parts = re.split(r"<#(\\d+(?:\\.\\d+)?)#>", script)
out = []
for i in range(0, len(parts), 2):
text = parts[i].strip()
pause = float(parts[i + 1]) if i + 1 < len(parts) else 0.0
if text:
out.append((text, pause))
return out
print(split_pauses("Welcome.<#1.5#>Today we launch.<#0.5#>Ready?"))Where does the silence go?
Sume's timeline audio concat is deliberately gapless, with no silence at the seams, and its parts accept only url, source_in and duration. So concat will not insert a gap. For audio you hand off to an editor, add the gaps there, using the pause list from the script above.
For video, Avatar videos allow a voice.type of silence, a non-speaking beat whose duration you set, and Timeline 1.0 has an audio.mode of silence for a declared length with no spine file. Both are video-side tools, not a way to make a standalone audio file with gaps.
If a precise gap inside a TTS audio file is a hard requirement, say so before you commit to Sume: this is a real gap, and MiniMax's markup solves it more directly.
Sources
Related posts
More in Comparisons
- MiniMax Video Agent template API: submit, poll, download
MiniMax Video Agent builds a video from a template_id plus your media and text. The endpoints, statuses, and what Sume offers for template-driven video.
- Mirage Tesseract free local engine vs a hosted avatar video API
Mirage Tesseract is an agent video suite with a free local engine. When does a hosted avatar API like Sume Avatar 1.0 fit better? Differences, with dated facts.
- Mirelo SFX video-to-sound vs how Sume adds sound to video
Mirelo generates sound effects synced to existing video. Sume has no video-to-SFX route; here is what it does offer for sound on a clip, and where each fits.
- Mubert Render 25-minute tracks and licence limits vs Sume Music
Mubert Render makes tracks up to 25 minutes on paid plans but bars streaming release. Sume Music makes song-length tracks at $0.125 each; loop for longer.
Written by Sume