WCAG 1.2.3 for an AI avatar video: is the script enough?
A talking-head avatar video already has its words in a script. Here is when that text meets WCAG 1.2.3 and what to add if the picture carries more.
The short answer
For a prerecorded avatar video where one avatar speaks a script to camera, WCAG 1.2.3 is usually met by publishing that script as a text alternative next to the video. The W3C Understanding page for 1.2.3 (read 2026-10-04) lists it as Level A, says a full text alternative satisfies it, and describes an exception when all the important information in the video track is already in the audio track. A talking head that only says what the script says is the easy case.
The catch is the word important. If the picture shows a product label, a price on a screen or a gesture the viewer needs, the script has to say it too, or you need a media alternative that does.
Why avatar video is a good fit
You write the words before the video exists. Sume Avatar 1.0 takes a script on POST /v1/avatar-1.0/talking-video and renders one avatar speaking it, so the transcript is not something you recover afterwards; it is the input. The Avatar videos docs describe the endpoint: provide exactly one of script or video_inputs, with the estimated duration inside 4 to 60 seconds.
That makes the text alternative a publishing step rather than a transcription job. Store the script string with the job, then render it on the page under the player.
When the script is not enough
Avatar clips can carry a product image or a scene, both optional fields in the same docs. Once the viewer must see something to follow the message, add the missing information in words.
- Product shown on screen: name it and state the detail that matters in the script itself.
- On-screen text in a scene: repeat it in the spoken script or in the text under the player.
- Silent beats: a clip with
voice.type: silencehas no speech to transcribe, so describe what happens (see silent avatar clips and WCAG 1.2.1). - Burned-in captions help hearing-loss viewers but are not a media alternative for blind viewers; they are a different success criterion.
What the W3C page does not say
The Understanding page is guidance, not a legal ruling, and it does not tell you how a particular regulator will treat generated video. It also says nothing about avatars specifically; the reasoning here is that an avatar clip is ordinary prerecorded video once it is rendered. If you publish under a rule that cites a conformance level, check which success criteria it names and keep the script text and a captioned version both available. Level A is the floor for 1.2.3, so there is no reason to skip it.
A quick decision table
Use this to decide what to publish under the player.
| Clip content | Script as text alternative enough? | Add |
|---|---|---|
| One avatar speaking to camera, nothing else important on screen | Yes | Nothing beyond the script text |
| Avatar plus a product image the viewer must see | Only if the script names the key details | Product details in the script or a short description |
| Avatar plus scene with readable text | Only if the text is also spoken | Repeat the text in the script |
| Silent segment with a visual action | No | Written description of the action |
Publish it next to the clip
Keep the script, the job id and the final MP4 URL together in your own records, then print the script in a labelled section below the player. Check it against the final video once, the same way you would check captions, because a script edited after rendering is no longer the alternative for that file. Sume mirrors rendered output to media.sume.com, so the video URL you store is stable for your page.
A short publishing routine
Do the same four steps for every avatar clip you embed, and the criterion stops being a per-video debate.
First, keep the exact script string from the request in your content store. Second, after the render completes, watch the clip once and confirm that nothing important appears only in the picture. Third, print the script under the player in a section with a plain heading, and link the heading from the player's label so keyboard and screen reader users can find it. Fourth, if you later re-render with a new script or a different product image, replace the published text at the same time.
This is cheap because nothing has to be transcribed: the words you submitted are the words that were spoken. It also pairs well with the first-frame previews, since a preview stills check is where you notice that a product image, rather than the speech, carries a detail you should add to the script.
Sources
Related posts
More in Sume Avatar 1.0
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
- Avatar Face Swap API (Beta): apply an avatar face to a video
Avatar Face Swap 1.0 is a Beta Sume endpoint that applies a ready avatar's face to a short public source video. Required fields, limits, and polling.
- How to create a reusable AI avatar with the Sume Avatar 1.0 API
Send POST /v1/avatar-1.0/generate with an avatar_handle and a prompt, profile, or image input. Poll the job, then reuse the handle for avatar videos.
- Multi-scene avatar video API: build one video from ordered scenes
Send ordered video_inputs instead of one script to compose spoken and silent scenes into one 4-60 second avatar video. Scene fields, rules, and limits.
Written by Sume