WCAG 1.2.5 audio description: narrate a video with TTS timestamps

WCAG 1.2.5 asks for audio description on prerecorded video. Build the narration track with Sume TTS word timings and timeline audio concat.

5 min readSume
All posts

WCAG success criterion 1.2.5 is a Level AA requirement: prerecorded video in synchronized media needs audio description, which is narration of the important visual information that the soundtrack does not already convey. If every visual fact is already spoken, W3C's page says the criterion can be met without it. For everything else, you need a narration track. Sume can produce the audio part: write the description, synthesize it with TTS, and use word timings to place each line.

What the W3C page says

The page is guidance, not Sume's reading of the law. Check the exact conformance level your contract or regulation names.

Criterion, read 2026-10-02
ItemDetail
Criterion1.2.5 Audio Description (Prerecorded)
LevelAA
Applies toPrerecorded video in synchronized media
ExceptionNot needed if all visual information is already in the audio
What it isNarration describing important visual details

Build the narration with Sume

Describing what is on screen is an editorial task: write short lines that fit the gaps. Then send each line to POST /v1/tts-1.0/generate with timestamps.words set to true. With segmentation.mode set to sentence and a wav output, each sentence comes back as its own slice with gapless start and end times.

For wav slices set the container to wav. mp3 returns timings only, without per-segment audio URLs.

import os, requests

r = requests.post(
    "https://api.sume.com/v1/tts-1.0/generate",
    headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
             "Idempotency-Key": "tts-demo-001"},
    json={"transcript": "Welcome back. Today we cover three updates.",
          "avatar_id": os.environ["SUME_AVATAR_ID"],
          "generation_config": {"speed": 0.95, "emotion": "warm"},
          "output_format": {"container": "wav", "encoding": "pcm_s16le"},
          "timestamps": {"words": True},
          "segmentation": {"mode": "sentence"}},
)
print(r.status_code, r.json())

Assemble and check

Join the lines into one narration file with timeline audio concat (1 to 20 parts per job), or place them in a Timeline 1.0 render. Then listen against the video: audio description is judged by whether a blind viewer gets the missing information, not by file validity.

Label the narration as synthetic where your policy requires it, and keep the caption track as a separate deliverable.

Worked example

A short worked pass helps. Suppose a 90-second product video has a 6-second silent section where a person opens a box.

  • Write a description that fits the gap, for example one 12-word sentence.
  • Synthesize it with sentence segmentation and wav output, then read each segment's start and end.
  • Concatenate the narration lines with the original dialogue track in timeline audio, or place the narration over the gap in a render.

Checklist before you commit

Remember that W3C's page also covers what counts as important visual information. When in doubt, ask a screen-reader user to listen, and keep the text version of your descriptions for people who prefer reading.

  • Confirm which conformance level applies to you.
  • Keep description lines short enough to fit pauses.
  • Ship the captions separately; audio description is not a substitute for them.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume