Silent avatar video clip: what WCAG 1.2.1 asks for
A silent beat or silent clip in an avatar video is prerecorded video-only content. Here is what WCAG 1.2.1 wants and how to supply it from Sume.
The answer
If an avatar clip has no speech at all, it is prerecorded video-only media, and WCAG 2.2 success criterion 1.2.1 applies: the W3C Understanding page (read 2026-10-04) asks for an alternative that presents equivalent information, either a text alternative or an audio track. The related non-text content criterion 1.1.1 (read 2026-10-04) says prerecorded video-only files need a text alternative, with descriptive identification as the minimum for time-based media.
This matters for Sume because voice.type: "silence" exists on the talking-video endpoint. Per the Avatar videos docs, a silence voice needs a duration and allows no script, so you end up with an avatar that moves without speaking.
What counts as the silent case
Not every quiet moment is a 1.2.1 problem. The difference is whether the whole file is silent or only one beat in a clip that otherwise speaks.
- Whole clip silent (
voice.type: silence): treat the file as video-only and publish an alternative. - One pause inside a spoken video: the file has audio, so captions and a transcript cover the speech; describe the pause only if it carries meaning.
- Looping background clip with no information: a short descriptive label is the lower bar, but see looping avatar clips and WCAG 2.2.2 for motion controls.
Supplying the alternative
Write a sentence or two that says what the viewer would learn by watching, and place it next to the player. Because Sume gives you the job inputs (duration, avatar handle, optional product image or scene), you already know what is on screen; the description is a restatement of your own request, not a guess about the output.
If the clip is an intro that sets a mood and conveys nothing else, say that plainly. Descriptive identification is enough for that case in the 1.1.1 text above.
Quick reference
Match the file to the fix.
| Avatar clip | Has speech | What to publish |
|---|---|---|
| Standard talking-video with script | Yes | Script text, captions |
voice.type: silence with product image | No | Text alternative describing the product and action |
Silence segment between spoken video_inputs | Yes overall | Captions and transcript; describe only meaningful pauses |
Do not rely on captions here
Sume's inline captions burn in the spoken script text, so a silent clip has no words to caption. For a silent clip you would use the standalone captions endpoint with authored cues, which the docs describe as the route for silent clips, but those are overlay cards for sighted viewers. They do not replace the text alternative for assistive technology; put the description in the page text.
A worked example
Say you render a 6-second avatar clip with voice.type: silence, a product image of a kettle and a prompt scene of a kitchen counter, for a landing page hero. Nothing is spoken, nothing is captioned. The alternative could read: a presenter in a kitchen turns toward a kettle and gestures at it, with no speech. That one sentence identifies the content, which is the minimum for time-based media in the 1.1.1 text, and it takes ten seconds to write because you chose every ingredient in the request.
If the same hero later gains a spoken line, switch to the silence beat inside spoken video_inputs pattern, publish the script, and drop the description unless the pause itself matters.
Sources
Related posts
More in Sume Avatar 1.0
- WCAG 1.2.3 for an AI avatar video: is the script enough?
A talking-head avatar video already has its words in a script. Here is when that text meets WCAG 1.2.3 and what to add if the picture carries more.
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
- Avatar Face Swap API (Beta): apply an avatar face to a video
Avatar Face Swap 1.0 is a Beta Sume endpoint that applies a ready avatar's face to a short public source video. Required fields, limits, and polling.
- Avatar video previews: approve the first frame before rendering
Create an avatar video preview to get first-frame stills, regenerate them if needed, then call generate-video on the preview id to render the final video.
Written by Sume