HeyGen pause tag vs a Sume silence scene in an avatar video
HeyGen supports break tags up to 5 s on professional voice clones. Sume has no inline pause tag documented; you add a silence scene inside video_inputs instead.

HeyGen now documents <break time="0.5s"/> pauses for professional voice clones, up to five seconds each. Sume's documented equivalent is not a tag in the script: you add a scene whose voice.type is silence and give it a duration.
Both give you a deliberate beat in a talking video. They differ in where the pause lives, how long it can be, and what else changes at the same moment.
What did HeyGen change for pauses?
HeyGen's September 2026 changelog entry says its text-to-speech endpoints, POST /v3/models/audio/tts and POST /v3/models/audio/tts/stream, support explicit pauses with professional voice clones. Pause values use the <break time="0.5s"/> format in seconds, with a maximum of five seconds per pause. The same entry says other SSML tags remain unsupported for instant voice clones.
So the pause rides inside the text you send, and it only applies to the professional clone path the entry describes.
How do I add a pause on Sume?
Sume avatar videos accept ordered video_inputs instead of a single script. A scene with voice.type: "silence" is a non-speaking beat. duration is required, and script and input_text are not allowed on it. Spoken scenes keep type: "text" with exactly one of script or input_text.
That makes the pause a first-class scene with its own id and background, not a character inside a sentence.
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: avatar-pause-001" \
-d '{"avatar_handle":"sume_clawra","aspect_ratio":"9:16","video_inputs":[
{"id":"hook","voice":{"type":"text","script":"Wait for it.","duration":2},
"background":{"type":"prompt","prompt":"Bright home office"}},
{"id":"beat","voice":{"type":"silence","duration":2},
"background":{"type":"prompt","prompt":"Bright home office"}},
{"id":"cta","voice":{"type":"text","script":"Now you know.","duration":3},
"background":{"type":"prompt","prompt":"Bright home office"}}]}'How do the two approaches compare?
The table lists only what each vendor's own page states.
| HeyGen professional voice | Sume avatar video | |
|---|---|---|
| Where the pause lives | Inside the text as a break tag | A separate scene in video_inputs |
| Unit | Seconds in the tag value | duration on the scene |
| Documented maximum | Five seconds per pause | Total video must estimate to 4-60 seconds |
| Other SSML | Unsupported for instant clones | Not documented for scripts |
What are Sume's limits on silence beats?
The docs give no per-beat cap. The binding rule is the total: all scenes together must land in the 4-60 second window, and inline captions are rejected if the estimate is above 60 seconds. Current execution also supports one resolved avatar per final video and expects scene backgrounds to resolve to one shared scene, so a silence scene should reuse the same background as its neighbours.
Sume does not document a way to bend a pause inside a sentence. If you need a half-second breath mid-line, split the line into two spoken scenes with a short silence between them.
Does a pause change the cost or the window?
A silence scene is part of the planned duration, so it counts toward the 4-60 second window and toward what the render is estimated at. Sume reserves the estimated amount at submit and captures it when the job completes. Check the estimate on a short test before you scale a request that contains several beats.
Captions follow the spoken script text, so a silent scene has nothing to caption. That is expected behaviour, not a failure.
Which should you use?
If your pipeline already generates audio with HeyGen's professional clone, the tag is the shorter path. If you are building on Sume, plan beats as scenes and preview the result first: Avatar video previews show one first-frame still per scene before you pay for the full render. A worked version of this pattern is in add a pause to an AI avatar video.
Sources
Related posts
More in Comparisons
- HeyGen Studio Templates API vs a reusable Sume avatar request
HeyGen can now create templates from videos with POST /v3/templates. Sume has no avatar-video template object; the request body plus a handle is the template.
- HeyGen Video API 5-15 s at 768p vs Sume avatar video 4-60 s
HeyGen's new Video API makes 5-15 second clips at 480p or 768p from a prompt. Sume's avatar video renders scripts of 4-60 seconds at 720p. Which fits your job?
- HeyGen scene-scoped edits vs fixing one scene in a Sume avatar
HeyGen's Video Agent can edit one scene and keep the rest. Sume has no scene-edit call: structural changes need a new preview; only quality changes late.
- HeyGen's 22 Spanish and 17 Arabic variants vs Sume's language field
HeyGen lists regional variants such as 22 Spanish and 17 Arabic. Sume's TTS takes one BCP-47 language string and does not publish a per-region variant list.
Written by Sume