WCAG 1.2.3 for an AI avatar video: is the script enough?

A talking-head avatar video already has its words in a script. Here is when that text meets WCAG 1.2.3 and what to add if the picture carries more.

5 min readSume
All posts

The short answer

For a prerecorded avatar video where one avatar speaks a script to camera, WCAG 1.2.3 is usually met by publishing that script as a text alternative next to the video. The W3C Understanding page for 1.2.3 (read 2026-10-04) lists it as Level A, says a full text alternative satisfies it, and describes an exception when all the important information in the video track is already in the audio track. A talking head that only says what the script says is the easy case.

The catch is the word important. If the picture shows a product label, a price on a screen or a gesture the viewer needs, the script has to say it too, or you need a media alternative that does.

Why avatar video is a good fit

You write the words before the video exists. Sume Avatar 1.0 takes a script on POST /v1/avatar-1.0/talking-video and renders one avatar speaking it, so the transcript is not something you recover afterwards; it is the input. The Avatar videos docs describe the endpoint: provide exactly one of script or video_inputs, with the estimated duration inside 4 to 60 seconds.

That makes the text alternative a publishing step rather than a transcription job. Store the script string with the job, then render it on the page under the player.

When the script is not enough

Avatar clips can carry a product image or a scene, both optional fields in the same docs. Once the viewer must see something to follow the message, add the missing information in words.

  • Product shown on screen: name it and state the detail that matters in the script itself.
  • On-screen text in a scene: repeat it in the spoken script or in the text under the player.
  • Silent beats: a clip with voice.type: silence has no speech to transcribe, so describe what happens (see silent avatar clips and WCAG 1.2.1).
  • Burned-in captions help hearing-loss viewers but are not a media alternative for blind viewers; they are a different success criterion.

What the W3C page does not say

The Understanding page is guidance, not a legal ruling, and it does not tell you how a particular regulator will treat generated video. It also says nothing about avatars specifically; the reasoning here is that an avatar clip is ordinary prerecorded video once it is rendered. If you publish under a rule that cites a conformance level, check which success criteria it names and keep the script text and a captioned version both available. Level A is the floor for 1.2.3, so there is no reason to skip it.

A quick decision table

Use this to decide what to publish under the player.

Which alternative fits which avatar clip, based on W3C 1.2.3 Understanding page (read 2026-10-04)
Clip contentScript as text alternative enough?Add
One avatar speaking to camera, nothing else important on screenYesNothing beyond the script text
Avatar plus a product image the viewer must seeOnly if the script names the key detailsProduct details in the script or a short description
Avatar plus scene with readable textOnly if the text is also spokenRepeat the text in the script
Silent segment with a visual actionNoWritten description of the action

Publish it next to the clip

Keep the script, the job id and the final MP4 URL together in your own records, then print the script in a labelled section below the player. Check it against the final video once, the same way you would check captions, because a script edited after rendering is no longer the alternative for that file. Sume mirrors rendered output to media.sume.com, so the video URL you store is stable for your page.

A short publishing routine

Do the same four steps for every avatar clip you embed, and the criterion stops being a per-video debate.

First, keep the exact script string from the request in your content store. Second, after the render completes, watch the clip once and confirm that nothing important appears only in the picture. Third, print the script under the player in a section with a plain heading, and link the heading from the player's label so keyboard and screen reader users can find it. Fourth, if you later re-render with a new script or a different product image, replace the published text at the same time.

This is cheap because nothing has to be transcribed: the words you submitted are the words that were spoken. It also pairs well with the first-frame previews, since a preview stills check is where you notice that a product image, rather than the speech, carries a detail you should add to the script.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume