Tavus enable_closed_captions vs captions burned into a Sume clip
Tavus closed captions are a flag shown live in the conversation page; Sume burns captions into the MP4 itself, with a style, a language hint and a 60 s ceiling.
On Tavus, closed captions are a boolean, enable_closed_captions, that you set when you create the conversation; captions then appear in the web interface at the conversation_url whenever you or the agent speaks. On Sume, captions are part of the video file: an avatar video accepts a captions object and burns the style into the final MP4.
The difference decides where the clip can go. Live captions exist only inside the conversation page. Burned captions travel with the file to any feed that autoplays muted.
What each side documents
Tavus's page documents one parameter and one delivery path: set enable_closed_captions to true and captions show during the call in the page at conversation_url. It does not state a default (Tavus docs: Closed Captions).
Sume's avatar video takes captions: { enabled, style, language }, plus optional font and script_text. The default style is slam; the others are punch, tiktok-green, korean-ad and a set of Hangul styles (Generate avatar video).
| Question | Tavus | Sume avatar video |
|---|---|---|
| How you turn them on | enable_closed_captions: true | captions.enabled: true |
| Where they appear | The conversation page at conversation_url | Burned into the final MP4 |
| Styling | Not documented on the page | slam (default), punch, tiktok-green, Hangul styles |
| Length limit | Not stated | Inline captions rejected above 60 s estimated duration |
| Extra charge | Not stated | No separate billed video-caption job for inline captions (per the docs) |
Failure behaviour on Sume
Caption stage failures are soft: the avatar job can still succeed with a clean primary video_url and captions.status=failed. Preview stills are never captioned. A Korean script with a Latin-only style such as slam is rejected with 400 caption_hangul_text_latin_style rather than rendered as missing glyphs, so pick a Hangul style for Korean speech.
Check captions.status in the result before publishing, and keep the clean video_url as the fallback you can caption again with the standalone Video captions model.
- For a recorded Tavus call you want to caption afterwards, the file has to be at a public HTTPS URL first; see the linked post on captioning a recorded call.
- Use
script_textwhen the spoken words differ from the on-screen words, such as a brand name spelled for the voice but written differently.
Choosing a style
Burned captions are a design choice as much as an accessibility one. A heavy style such as slam is built for short vertical clips with a few words on screen at once. For longer explainers a quieter style reads better. Test the style on a real script at the final aspect ratio before you commit a batch.
If your audience is Korean, use a Hangul style and check a rendered sample, since the service refuses a Latin-only style on Hangul text. If the platform you post to adds its own captions, decide which set viewers should see so the two do not overlap.
What to verify before publishing
Play the finished file with the sound off and read it end to end. Check that line breaks do not split names, that numbers are written the way you want, and that nothing sits under a platform's own interface elements at the bottom of the frame.
If captions.status is failed, publish the clean video_url only if the platform adds its own captions, or caption it again with the standalone model before posting.
Sources
Related posts
More in Comparisons
- Tavus guardrails: 1,000-character rules vs reviewing a Sume script
Tavus guardrails steer a live agent with rules up to 1,000 characters, 50 per PAL, with no guarantee. A Sume avatar script can be reviewed before it renders.
- Tavus knowledge base limits vs putting the facts in a Sume script
Tavus knowledge bases are English-only, take 5-10 minutes to process and cap crawls at 10,000 documents. A Sume avatar script carries its facts in the text.
- Tavus Magic Canvas cards vs what a rendered Sume avatar clip shows
Tavus Magic Canvas shows 8 kinds of interactive cards in live video calls only. A Sume avatar clip is a fixed MP4: CTA goes in a closing scene.
- Tavus max_call_duration is plan-capped; Sume's cap is 60 s a job
Tavus ends a call at max_call_duration, capped by your plan; an unjoined call times out after 300 s. Sume's avatar video takes 4-60 s per job. Table inside.
Written by Sume