Avatar preview captions: stored at create, burned at generate-video

Preview stills are never captioned. Caption settings on the preview are stored, then applied at generate-video. What that means.

5 min readSume
All posts

If you add captions to an avatar video preview, are the preview stills captioned? No. Preview stills are never caption-burned. The caption settings you send on create are stored with the preview and applied only when you call generate-video on it.

That is useful now that voices change often. ElevenLabs released Eleven v4 on 2026-09-28 (read 2026-10-04) and other vendors shipped new voices the same month. Keeping caption intent separate from the voice and the framing means you can approve a still once and still change what is spoken.

The flow

The preview family has four routes: POST /v1/avatar-video-previews, GET /v1/avatar-video-previews/:id, POST /v1/avatar-video-previews/:id/regenerate and POST /v1/avatar-video-previews/:id/generate-video. Create takes the same body as an avatar video: exactly one of script or video_inputs, plus optional product_image, scene, quality, aspect_ratio, title and captions.

Inline captions use the same four knobs as standalone captions: style, an optional font, a language hint and script_text. Styles are slam (default), punch, tiktok-green, korean-ad, plus the Hangul identities, per the preview guide.

What happens where

Read the table as a checklist before you pay for the full render.

Where captions apply in the avatar preview flow, read 2026-10-04
StepCaptions
Create previewStored for later; the still is never captioned
Regenerate stillStill has no captions; stored intent unchanged
generate-videoCaptions burned into the clean final MP4
Failure at the caption stageThe avatar job can still succeed with a clean primary video and captions.status=failed

Limits to know

Estimated duration above 60 seconds is rejected for inline captions. A Korean script with slam, punch or tiktok-green is rejected with 400 caption_hangul_text_latin_style instead of being re-styled, because those faces render Hangul as tofu. Choose a Hangul style for Korean speech.

Because a caption failure soft-fails, check captions.status on the result before you ship, not only the job status. If it failed, you still hold a good clean video and can run a standalone video captions job on it.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume