ADA Title II April 26, 2027: caption course videos with an API

ADA.gov sets WCAG 2.1 AA with a 2027 deadline for larger public bodies. See how to burn captions into lecture clips with Sume and what it leaves undone.

5 min readSume
All posts

A public college or school district that falls under the ADA Title II web rule needs captions on prerecorded lecture video by April 26, 2027 if it serves 50,000 or more people, and by April 26, 2028 if it serves fewer. Sume can burn captions into a clip through POST /v1/video-captions for $0.20 per video up to 60 seconds, but that covers one piece of the work: it does not export caption files, write audio description or certify compliance.

The dates and the standard come from ADA.gov's web rule first steps page and the caption rule from W3C's Understanding Success Criterion 1.2.2, both read on 2026-10-02. Sume's side comes from Video captions. This is a tooling post, not legal advice: ask your counsel which rule and deadline apply to your institution.

What does the ADA Title II web rule require, and by when?

The ADA.gov page names Web Content Accessibility Guidelines version 2.1 Level AA as the required standard and says it is mandatory even though the title says guidelines. It lists one date per entity size, and the first date is the one most universities will care about.

The ADA.gov page does not itself discuss captions or video. For that you read the WCAG success criterion, which is where the caption requirement lives.

Title II compliance dates on ADA.gov (read 2026-10-02)
EntityCompliance deadline
Population of 50,000 or moreApril 26, 2027
Fewer than 50,000 residentsApril 26, 2028
Special district governmentsApril 26, 2028

What does WCAG say about captions on recorded lectures?

Success Criterion 1.2.2 is a Level A criterion: captions must be provided for all prerecorded audio content in synchronized media, unless the media is a clearly labeled alternative to text. A Level AA target sits on top of Level A, so a lecture video with speech needs captions either way.

W3C's explanation is wider than a transcript of speech. Captions should carry dialogue, speaker identification and location, and sound effects, music or laughter that a viewer needs to follow the content. An automatic speech transcript usually covers only the first of those.

  • Dialogue: what is said, in sync with the picture.
  • Speaker identification: who is talking when it is not obvious on screen.
  • Non-speech audio: a lab buzzer, applause, a recorded clip played in class.

How do you caption a lecture clip with Sume?

Standalone video captions take a public HTTPS video URL and return a captioned video as a job. Speech-to-text runs when you omit text, and you can pass script_text when you already have the approved wording; Sume keeps the speech timings and aligns your script to them. Pass language as a hint, or omit it for automatic detection.

The call below is the whole request. It needs an Idempotency-Key header on retries and returns a job you poll with GET /v1/jobs/{id}/status, or you can ask for a webhook.

Batch the work by course. Name each caption job with the course and week in its Idempotency-Key, so a retry after a timeout reuses the job and the inventory sheet can map each finished video back to its lecture.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: lecture-04-captions-001" \
  -d '{
    "video_url": "https://example.com/lectures/week-04-clip.mp4",
    "language": "en",
    "style": "black-outline",
    "script_text": "Today we cover the central limit theorem."
  }'

Where does a burned-in caption fall short for accessibility?

Sume's captions are burned into the picture. The docs say SRT uploads are unsupported, and the public result carries the captioned video, not a transcript or caption file. That means viewers cannot switch them off, resize them in the player or use them with a screen reader. Many institutions still publish a sidecar caption track and a transcript beside the video; this API does not produce either.

Two limits are worth planning around. First, a caption job reserves $0.20 for videos up to 60 seconds under the current fixed estimate, and the docs say to confirm live pricing in the catalog, so a 50-minute lecture is a poor fit for one call: cut it with video trim and caption the pieces, as covered in captioning a long video. Second, speech-to-text is a draft. Names, formulas and course terms need your script_text or a human pass.

What should an institution do before the deadline?

Treat captions as one line in an inventory. List every prerecorded video, mark which have accurate captions, and decide which need a sidecar file as well as an open caption. Use video inspect with transcribe: true to get a transcript with word timings for the review pass, then feed corrected text back as script_text.

For sound cues and speaker labels, author them yourself as cues with text, start and end, which is the same path Sume uses for silent clips; the checklist in sound cues for deaf viewers walks through it. Audio description for visual-only content is a different requirement that Sume does not generate.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume