Tavus enable_closed_captions vs captions burned into a Sume clip

Tavus closed captions are a flag shown live in the conversation page; Sume burns captions into the MP4 itself, with a style, a language hint and a 60 s ceiling.

5 min readSume
All posts

On Tavus, closed captions are a boolean, enable_closed_captions, that you set when you create the conversation; captions then appear in the web interface at the conversation_url whenever you or the agent speaks. On Sume, captions are part of the video file: an avatar video accepts a captions object and burns the style into the final MP4.

The difference decides where the clip can go. Live captions exist only inside the conversation page. Burned captions travel with the file to any feed that autoplays muted.

What each side documents

Tavus's page documents one parameter and one delivery path: set enable_closed_captions to true and captions show during the call in the page at conversation_url. It does not state a default (Tavus docs: Closed Captions).

Sume's avatar video takes captions: { enabled, style, language }, plus optional font and script_text. The default style is slam; the others are punch, tiktok-green, korean-ad and a set of Hangul styles (Generate avatar video).

Live captions on Tavus and burned captions on Sume (read 2026-10-03)
QuestionTavusSume avatar video
How you turn them onenable_closed_captions: truecaptions.enabled: true
Where they appearThe conversation page at conversation_urlBurned into the final MP4
StylingNot documented on the pageslam (default), punch, tiktok-green, Hangul styles
Length limitNot statedInline captions rejected above 60 s estimated duration
Extra chargeNot statedNo separate billed video-caption job for inline captions (per the docs)

Failure behaviour on Sume

Caption stage failures are soft: the avatar job can still succeed with a clean primary video_url and captions.status=failed. Preview stills are never captioned. A Korean script with a Latin-only style such as slam is rejected with 400 caption_hangul_text_latin_style rather than rendered as missing glyphs, so pick a Hangul style for Korean speech.

Check captions.status in the result before publishing, and keep the clean video_url as the fallback you can caption again with the standalone Video captions model.

  • For a recorded Tavus call you want to caption afterwards, the file has to be at a public HTTPS URL first; see the linked post on captioning a recorded call.
  • Use script_text when the spoken words differ from the on-screen words, such as a brand name spelled for the voice but written differently.

Choosing a style

Burned captions are a design choice as much as an accessibility one. A heavy style such as slam is built for short vertical clips with a few words on screen at once. For longer explainers a quieter style reads better. Test the style on a real script at the final aspect ratio before you commit a batch.

If your audience is Korean, use a Hangul style and check a rendered sample, since the service refuses a Latin-only style on Hangul text. If the platform you post to adds its own captions, decide which set viewers should see so the two do not overlap.

What to verify before publishing

Play the finished file with the sound off and read it end to end. Check that line breaks do not split names, that numbers are written the way you want, and that nothing sits under a platform's own interface elements at the bottom of the frame.

If captions.status is failed, publish the clean video_url only if the platform adds its own captions, or caption it again with the standalone model before posting.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume