MiniMax <#1.5#> pause markup vs Sume TTS: no pause marker

MiniMax inserts pauses with <#x#> markers from 0.01 to 99.99 seconds. Sume TTS documents no pause marker; here is how to split a script and what to use instead.

5 min readSume
All posts

Can I add a timed pause in Sume TTS the way MiniMax does?

Not with a documented marker. MiniMax lets you write a pause inside the text as <#x#>, where x is the duration in seconds from 0.01 to 99.99 with up to two decimals. The Sume TTS 1.0 request has no pause field and documents no marker syntax, so a script containing <#1.5#> is not a supported way to get a 1.5-second gap.

You can still shape pacing on Sume. The sections below split the work into what the text can do, what segmentation gives you and where an edit step is needed.

What does the Sume contract give you for pacing?

Three things. Punctuation and paragraph breaks in the transcript, which the voice reads as natural pauses. generation_config.speed from 0.6 to 1.5 to slow the whole read. And segmentation: with timestamps.words and segmentation.mode set to sentence, you receive gapless sentence segments, where boundary_lead_ms (0 to 500, default 70) sets how long after the last word each cut falls and the next segment absorbs the pause.

That last control moves a cut point; it does not synthesize silence. Treat it as a way to keep natural breath gaps inside the slices, not as a pause generator.

Pause tools, MiniMax reference and Sume OpenAPI (read 2026-10-02)
NeedMiniMax T2ASume TTS 1.0
Timed pause in the text<#x#>, 0.01 to 99.99 sNo documented marker
Slower read overallspeed 0.5 to 2generation_config.speed 0.6 to 1.5
Where sentences cutNot covered heresegmentation.boundary_lead_ms 0 to 500, default 70
Silent beat in videoNot covered hereAvatar video voice.type silence, duration required

How do you port a script that already uses <#x#> markers?

Split at each marker, generate each stretch as its own job and keep the pause lengths in a list. The helper below returns pairs of text and the pause that followed it, so your edit step knows where gaps belong.

import re

def split_pauses(script):
    parts = re.split(r"<#(\\d+(?:\\.\\d+)?)#>", script)
    out = []
    for i in range(0, len(parts), 2):
        text = parts[i].strip()
        pause = float(parts[i + 1]) if i + 1 < len(parts) else 0.0
        if text:
            out.append((text, pause))
    return out

print(split_pauses("Welcome.<#1.5#>Today we launch.<#0.5#>Ready?"))

Where does the silence go?

Sume's timeline audio concat is deliberately gapless, with no silence at the seams, and its parts accept only url, source_in and duration. So concat will not insert a gap. For audio you hand off to an editor, add the gaps there, using the pause list from the script above.

For video, Avatar videos allow a voice.type of silence, a non-speaking beat whose duration you set, and Timeline 1.0 has an audio.mode of silence for a declared length with no spine file. Both are video-side tools, not a way to make a standalone audio file with gaps.

If a precise gap inside a TTS audio file is a hard requirement, say so before you commit to Sume: this is a real gap, and MiniMax's markup solves it more directly.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume