Avatar preview captions: stored at create, burned at generate-video
Preview stills are never captioned. Caption settings on the preview are stored, then applied at generate-video. What that means.
If you add captions to an avatar video preview, are the preview stills captioned? No. Preview stills are never caption-burned. The caption settings you send on create are stored with the preview and applied only when you call generate-video on it.
That is useful now that voices change often. ElevenLabs released Eleven v4 on 2026-09-28 (read 2026-10-04) and other vendors shipped new voices the same month. Keeping caption intent separate from the voice and the framing means you can approve a still once and still change what is spoken.
The flow
The preview family has four routes: POST /v1/avatar-video-previews, GET /v1/avatar-video-previews/:id, POST /v1/avatar-video-previews/:id/regenerate and POST /v1/avatar-video-previews/:id/generate-video. Create takes the same body as an avatar video: exactly one of script or video_inputs, plus optional product_image, scene, quality, aspect_ratio, title and captions.
Inline captions use the same four knobs as standalone captions: style, an optional font, a language hint and script_text. Styles are slam (default), punch, tiktok-green, korean-ad, plus the Hangul identities, per the preview guide.
What happens where
Read the table as a checklist before you pay for the full render.
| Step | Captions |
|---|---|
| Create preview | Stored for later; the still is never captioned |
| Regenerate still | Still has no captions; stored intent unchanged |
| generate-video | Captions burned into the clean final MP4 |
| Failure at the caption stage | The avatar job can still succeed with a clean primary video and captions.status=failed |
Limits to know
Estimated duration above 60 seconds is rejected for inline captions. A Korean script with slam, punch or tiktok-green is rejected with 400 caption_hangul_text_latin_style instead of being re-styled, because those faces render Hangul as tofu. Choose a Hangul style for Korean speech.
Because a caption failure soft-fails, check captions.status on the result before you ship, not only the job status. If it failed, you still hold a good clean video and can run a standalone video captions job on it.
Sources
Related posts
More in Sume Avatar 1.0
- Avatar video script too long? The 4 to 60 second rule
Sume Avatar 1.0 accepts scripts it estimates at 4 to 60 seconds. Estimate yours first, then shorten it or split it into separate videos before you submit.
- Avatar validation failed: fix it, or skip verification
Searches for HeyGen avatar validation failed show people stuck on verification. On Sume an avatar create is a job you poll, and failures return a job error.
- Crowdfunding intro video: preview first, render at Max
Check a 30-second campaign intro on a Sume preview, then override quality to max only at generate-video. A 30-second Max render costs $16.50.
- Denmark's likeness bill: 50 years after death, avatar checks
Denmark's likeness bill proposes protection for 50 years after death, and the European Commission has raised concerns. What to check before avatar videos.
Written by Sume