Live captions vs rendered avatar clip captions: WCAG 1.2.4
Full-duplex video AI is live, so WCAG 1.2.4 applies. A rendered avatar clip is prerecorded and follows 1.2.2 instead. Where Sume captions fit.
Which rule applies
Pre-rendered avatar video falls under WCAG 1.2.2, captions for prerecorded media; live conversational video falls under 1.2.4, captions for live media. Tavus announced Griffin on 2026-10-01 as a full-duplex video-to-video system, which puts that kind of product on the live side of the line (read 2026-10-04). Sume Avatar 1.0 is the other side: you submit a job and receive an MP4.
The 1.2.4 Understanding page (read 2026-10-04) lists the criterion as Level AA and says computer-generated content does not count as live for this purpose, while the 1.2.2 page (read 2026-10-04) lists Level A, accepts open captions as meeting it and also covers speaker identification and non-speech information.
What Sume produces
Sume does not run a live avatar session. The talking-video endpoint renders one avatar speaking a script into a file, and offers inline captions with enabled, style and language. They burn into the final MP4 from the script text. For a clip you already have, POST /v1/video-captions burns styled captions after the fact, as the Video captions docs describe.
Because the output is a prerecorded file, the 1.2.2 route is available: ship the clip with burned-in captions and a transcript. A live product has no such luxury, because its captions must be produced while the conversation happens.
Side by side
The decision is mostly about timing.
| Question | Rendered avatar clip (Sume) | Live conversational video |
|---|---|---|
| WCAG criterion | 1.2.2 Captions (Prerecorded), Level A | 1.2.4 Captions (Live), Level AA |
| When are captions made | Before publishing, burned in at render | During the session |
| Can you review captions first | Yes | No |
| Delivery | MP4 with optional inline captions | Streaming session |
What to do if you need both
Teams that pair a live agent with recorded explainers meet two criteria. For the recorded part, set captions.enabled on the talking-video request. If the caption stage fails, the docs say the job can still succeed with a clean video and captions.status=failed, so check that field before publishing and re-caption the clean video with the standalone endpoint, as described in the soft-fail post.
For the live part, captioning is a property of your live stack, not of Sume. Read what Griffin changes for how the two product types divide the work.
The practical takeaway
Do not pick the criterion by how the video looks; pick it by when the audio was produced. If you wrote the words and rendered them, you can caption, check and publish. If the words appear only as someone speaks, you need a live captioning source. Sume covers the first case, and says so in its docs.
Sources
Related posts
More in Comparisons
- xAI batch mixes chat, image and video in one file; Sume uses queues
xAI's Batch API now takes chat, image and video requests in one JSONL file. A Sume bulk queue targets exactly one Format. How to split a holiday job.
- xAI batch video URLs expire in 1 hour: size the download worker
xAI's release notes say image and video URLs in batch results expire after 1 hour. How many parallel downloads you need, and how Sume's durable URLs differ.
- xAI file URLs auto-expire; Sume fetches image_url once at create
xAI's Files API can serve public URLs that auto-expire. Sume fetches an attachment image_url at run creation and copies it. What that means for hosting.
- Sume vs Argil: AI avatar video and video agents compared
Argil makes AI-avatar and story videos with a chat agent, Director; Sume is a video agent with a multi-model API. Avatars, API, pricing, and limits compared.
Written by Sume