HeyGen lipsync captions are always on; Sume's are opt-in
HeyGen deprecated enable_caption on translation and lipsync and now always returns SRT and VTT. Sume avatar videos burn captions only when you ask for them.

On HeyGen, you no longer choose captions for a lipsync or video translation job: the caption flags are deprecated and ignored, and the completed response always carries caption URLs. On Sume, it is the opposite: an avatar video has no captions unless you send a captions object, and then they are burned into the MP4.
The HeyGen detail comes from its developer changelog, where an August 2026 entry says caption-generation flags on video translation and lipsync requests are deprecated and ignored, enable_caption is deprecated across several endpoints, and completed responses always expose caption URLs in SRT and VTT. The rest of this post is about what that means if you are moving a pipeline between the two.
What changed in HeyGen's lipsync request?
Sending enable_caption no longer changes anything. Captions are generated for every translation and lipsync, and you get a sidecar file (subtitle_url) whether or not you wanted one. You decide at display time: show the SRT or VTT in your player, download it, or ignore it.
The changelog does not say whether a sidecar is billed separately, so check HeyGen's pricing page before assuming it is free.
How do captions work on a Sume avatar video?
Captions are an optional field on the talking-video request. captions burns a style into the clean final MP4 after generation, using the script or video_inputs text, and preview stills are never captioned. The default style is slam; others include punch, tiktok-green and korean-ad.
The result is a video with the words in the pixels, not a sidecar. If you need a separate SRT or VTT for a player, a burned-in file is the wrong output; Sume's page describes burned-in styles only.
| Question | HeyGen lipsync and translation | Sume avatar video |
|---|---|---|
| Do I choose captions at submit? | No, flags are ignored | Yes, add a captions object |
| Default result | Caption URLs always in the response | Clean MP4, no captions |
| Caption format | SRT and VTT | Burned into the MP4 |
| Billing for captions | Not stated on the changelog | Inline captions are not a separate billed caption job |
| Existing public video | Not covered in the changelog | Standalone Video captions burns captions onto a public video URL |
What fails differently on Sume?
Inline captions have two rules that a sidecar does not. A script estimated above 60 seconds is rejected for inline captions, and a Korean script with slam, punch or tiktok-green is rejected with 400 caption_hangul_text_latin_style, because those faces render Hangul as tofu. Use a Hangul style for Korean speech.
A caption failure does not lose the video. Caption stage failures soft-fail: the job can still succeed with a clean video_url and captions.status=failed, and you can caption that clean file later through the standalone Video captions endpoint.
Which should I pick for a multi-market pipeline?
Pick on the output you need, not on the vendor. If your player wants selectable text, language switching or accessibility tracks, an always-returned SRT or VTT is the better shape, and you should treat it as the contract on HeyGen. If the clip goes straight to a feed where the platform strips sidecars, burned-in captions are what you need, and on Sume you opt in per request.
Whichever side you are on, test the cases the changelog does not spell out: a silent clip, a Korean script, and a script over 60 seconds.
What does a migration checklist look like?
Moving a lipsync pipeline from HeyGen to Sume, delete the code that sets caption flags and decide where captions come from. If your downstream step reads subtitle_url, replace it: on Sume there is no sidecar to read, so either request burned-in captions on the avatar job or run the standalone Video captions step on the finished file.
Going the other way, stop passing enable_caption and handle the case where a sidecar exists that you never asked for. In both directions, add a test for a clip whose speech is not in the language of the caption style, since style choice and language hints are where the two behave least alike.
Sources
Related posts
More in Sume Avatar 1.0
- Holiday avatar ad roster: 3 presenters x 4 scripts, cost by tier
Three reusable avatars and four holiday scripts make 12 clips. Priced on standard, plus and max with a product image; the avatars cost $2.85.
- LemonSlice API: image plus streaming audio vs Sume lip sync
LemonSlice drives a live avatar from an image and streaming audio. For a recorded line, Sume's lip sync takes a still plus an audio_url and returns a file.
- Lip sync audio under 5 seconds: H3 Max rejects it, use Fabric
Sume's MiniMax H3 Max lip sync only accepts 5 to 14.8 seconds of audio and returns invalid_request outside it. A 4.5-second line goes to Fabric instead.
- Lip sync looks fake? Check teeth, profile, timing and seams on Sume
sync. labs names four tells of fake lip sync. Pull PNG stills from a Sume avatar clip with video frames at the moments each tell shows up, then decide.
Written by Sume