Silent avatar video clip: what WCAG 1.2.1 asks for

A silent beat or silent clip in an avatar video is prerecorded video-only content. Here is what WCAG 1.2.1 wants and how to supply it from Sume.

5 min readSume
All posts

The answer

If an avatar clip has no speech at all, it is prerecorded video-only media, and WCAG 2.2 success criterion 1.2.1 applies: the W3C Understanding page (read 2026-10-04) asks for an alternative that presents equivalent information, either a text alternative or an audio track. The related non-text content criterion 1.1.1 (read 2026-10-04) says prerecorded video-only files need a text alternative, with descriptive identification as the minimum for time-based media.

This matters for Sume because voice.type: "silence" exists on the talking-video endpoint. Per the Avatar videos docs, a silence voice needs a duration and allows no script, so you end up with an avatar that moves without speaking.

What counts as the silent case

Not every quiet moment is a 1.2.1 problem. The difference is whether the whole file is silent or only one beat in a clip that otherwise speaks.

  • Whole clip silent (voice.type: silence): treat the file as video-only and publish an alternative.
  • One pause inside a spoken video: the file has audio, so captions and a transcript cover the speech; describe the pause only if it carries meaning.
  • Looping background clip with no information: a short descriptive label is the lower bar, but see looping avatar clips and WCAG 2.2.2 for motion controls.

Supplying the alternative

Write a sentence or two that says what the viewer would learn by watching, and place it next to the player. Because Sume gives you the job inputs (duration, avatar handle, optional product image or scene), you already know what is on screen; the description is a restatement of your own request, not a guess about the output.

If the clip is an intro that sets a mood and conveys nothing else, say that plainly. Descriptive identification is enough for that case in the 1.1.1 text above.

Quick reference

Match the file to the fix.

Alternative needed by clip type, from W3C 1.2.1 and 1.1.1 Understanding pages (read 2026-10-04)
Avatar clipHas speechWhat to publish
Standard talking-video with scriptYesScript text, captions
voice.type: silence with product imageNoText alternative describing the product and action
Silence segment between spoken video_inputsYes overallCaptions and transcript; describe only meaningful pauses

Do not rely on captions here

Sume's inline captions burn in the spoken script text, so a silent clip has no words to caption. For a silent clip you would use the standalone captions endpoint with authored cues, which the docs describe as the route for silent clips, but those are overlay cards for sighted viewers. They do not replace the text alternative for assistive technology; put the description in the page text.

A worked example

Say you render a 6-second avatar clip with voice.type: silence, a product image of a kettle and a prompt scene of a kitchen counter, for a landing page hero. Nothing is spoken, nothing is captioned. The alternative could read: a presenter in a kitchen turns toward a kettle and gestures at it, with no speech. That one sentence identifies the content, which is the minimum for time-based media in the 1.1.1 text, and it takes ten seconds to write because you chose every ingredient in the request.

If the same hero later gains a spoken line, switch to the silence beat inside spoken video_inputs pattern, publish the script, and drop the description unless the pause itself matters.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume