HeyGen pause tag vs a Sume silence scene in an avatar video

HeyGen supports break tags up to 5 s on professional voice clones. Sume has no inline pause tag documented; you add a silence scene inside video_inputs instead.

4 min readSume
All posts

HeyGen now documents <break time="0.5s"/> pauses for professional voice clones, up to five seconds each. Sume's documented equivalent is not a tag in the script: you add a scene whose voice.type is silence and give it a duration.

Both give you a deliberate beat in a talking video. They differ in where the pause lives, how long it can be, and what else changes at the same moment.

What did HeyGen change for pauses?

HeyGen's September 2026 changelog entry says its text-to-speech endpoints, POST /v3/models/audio/tts and POST /v3/models/audio/tts/stream, support explicit pauses with professional voice clones. Pause values use the <break time="0.5s"/> format in seconds, with a maximum of five seconds per pause. The same entry says other SSML tags remain unsupported for instant voice clones.

So the pause rides inside the text you send, and it only applies to the professional clone path the entry describes.

How do I add a pause on Sume?

Sume avatar videos accept ordered video_inputs instead of a single script. A scene with voice.type: "silence" is a non-speaking beat. duration is required, and script and input_text are not allowed on it. Spoken scenes keep type: "text" with exactly one of script or input_text.

That makes the pause a first-class scene with its own id and background, not a character inside a sentence.

curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: avatar-pause-001" \
  -d '{"avatar_handle":"sume_clawra","aspect_ratio":"9:16","video_inputs":[
    {"id":"hook","voice":{"type":"text","script":"Wait for it.","duration":2},
     "background":{"type":"prompt","prompt":"Bright home office"}},
    {"id":"beat","voice":{"type":"silence","duration":2},
     "background":{"type":"prompt","prompt":"Bright home office"}},
    {"id":"cta","voice":{"type":"text","script":"Now you know.","duration":3},
     "background":{"type":"prompt","prompt":"Bright home office"}}]}'

How do the two approaches compare?

The table lists only what each vendor's own page states.

Pause controls compared (HeyGen changelog and Sume docs, read 2026-10-02)
HeyGen professional voiceSume avatar video
Where the pause livesInside the text as a break tagA separate scene in video_inputs
UnitSeconds in the tag valueduration on the scene
Documented maximumFive seconds per pauseTotal video must estimate to 4-60 seconds
Other SSMLUnsupported for instant clonesNot documented for scripts

What are Sume's limits on silence beats?

The docs give no per-beat cap. The binding rule is the total: all scenes together must land in the 4-60 second window, and inline captions are rejected if the estimate is above 60 seconds. Current execution also supports one resolved avatar per final video and expects scene backgrounds to resolve to one shared scene, so a silence scene should reuse the same background as its neighbours.

Sume does not document a way to bend a pause inside a sentence. If you need a half-second breath mid-line, split the line into two spoken scenes with a short silence between them.

Does a pause change the cost or the window?

A silence scene is part of the planned duration, so it counts toward the 4-60 second window and toward what the render is estimated at. Sume reserves the estimated amount at submit and captures it when the job completes. Check the estimate on a short test before you scale a request that contains several beats.

Captions follow the spoken script text, so a silent scene has nothing to caption. That is expected behaviour, not a failure.

Which should you use?

If your pipeline already generates audio with HeyGen's professional clone, the tag is the shorter path. If you are building on Sume, plan beats as scenes and preview the result first: Avatar video previews show one first-frame still per scene before you pay for the full render. A worked version of this pattern is in add a pause to an AI avatar video.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume