Captions with sound cues for deaf viewers: W3C checklist on Sume
W3C says captions carry speech and non-speech sound. Sume's STT burn covers speech only, so here is how to author the full cue list and burn it.

Burned-in captions from Sume cover what is said, not what is heard. The W3C describes captions as a text version of "the speech and non-speech audio information needed to understand the content", so a creator who wants accessible captions has to add the sound cues themselves and burn the complete list as authored cues with Video captions.
This post walks through that workflow: transcribe once, merge in sound cues, burn, and review. It does not claim that Sume output meets any legal standard; that judgement stays with you.
What does the W3C say a caption must contain?
The W3C WAI captions page is the reference for what follows. These are the parts that matter for a small team:
- Speech and non-speech audio: the W3C definition covers both. A door slam or a phone ring is our own example of non-speech audio.
- Speaker identification: the page shows speaker labels in WebVTT, so label speakers where the viewer cannot tell who is talking.
- Synchronisation: captions are timed to the audio.
- Accuracy: the page says automatically generated captions do not meet accessibility requirements unless confirmed fully accurate, and usually need significant editing.
What does Sume's speech-to-captions path give you?
When you send only video_url (or script_text), Sume runs speech-to-text and burns the words it hears. That path needs audible speech: a silent clip fails as caption_no_speech. It produces no sound cues and no speaker names, because the transcript is words and timings only.
The other path is authored overlay copy. You pass cues (or segments), each with text, start and end in seconds, and Sume burns exactly that text at those times with no speech-to-text. The fields script_text, words, cues and segments are mutually exclusive, so a cue list has to hold everything you want on screen, speech included.
How do you build a complete cue list?
Get speech timings from Video inspect. It needs a media.sume.com artifact or asset of your workspace (import other clips first with POST /v1/media-imports) and an Idempotency-Key. Set transcribe: true with segmentation.mode: "sentence" and it returns gapless sentence segments[], shaped like caption lines. Transcription is billed at $0.01 per audio minute; probe and stills are not. Then:
1. Copy each sentence segment into a cue with its start and end.
2. Insert sound cues in square brackets, such as [door slams], in the gaps where the sound happens.
3. Prefix a speaker name when the picture does not make it obvious, such as MAYA:.
4. Correct the words by ear. Sume does not check your edits against the audio.
The docs describe no separate treatment for bracketed text, so sound cues burn in the same style as speech. Pick one style that stays legible for both.
| Element | Where it comes from | Burned by Sume? |
|---|---|---|
| Spoken words | Video inspect sentence segments, edited by you | Yes, as cues |
| Sound cues | You write them | Yes, as cues |
| Speaker names | You write them | Yes, as cue text |
| Sidecar caption file | Not produced | No: SRT uploads are unsupported |
What does the burn call look like?
The clip must be a fetchable public HTTPS video URL. Each accepted standalone job is $0.20 for videos up to 60 seconds under the current fixed estimate; check live pricing in GET /v1/catalog.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: access-captions-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/example/clean.mp4",
"style": "black-outline",
"cues": [
{"text": "[door slams]", "start": 0.0, "end": 1.2},
{"text": "MAYA: Welcome back to the studio.", "start": 1.4, "end": 3.6}
]
}'How should you review the result?
Play the burned clip with the sound on and then with the sound off. With sound on, check that each word matches what you hear and that each bracketed cue appears when the sound does. With sound off, check that the story still makes sense: if a viewer would miss a plot point, a cue is missing.
Keep cues short. A bracketed note such as [phone rings] reads in a glance, while a sentence of description pulls the eye away from the picture. If a cue needs fixing after the burn, the speech timings stay yours: change the cue list and run the caption job again, or restyle with source_caption_id when you only want a new look and not new words.
What does the whole pipeline cost?
For a 45-second clip the published rates are $0.01 for one minute of transcript through Video inspect and $0.20 for the burn, so about $0.21 per pass. Review time is the real cost, and it is the part that cannot be automated away. Rates can change; confirm them in GET /v1/catalog before a batch.
What does Sume not do here?
Sume burns captions into the picture, so viewers cannot turn them off, and it does not produce a caption file for a platform's own caption track. It does not detect non-speech sounds or identify speakers for you. If a platform needs a sidecar track, that is a separate tool in your pipeline; see burned-in versus sidecar captions for the trade-off.
Sources
Related posts
More in Media tools
- caption_no_speech on a silent clip: burn Halloween text with cues
A silent AI clip fails video captions with caption_no_speech. Pass cues with text, start and end instead: a worked giveaway announcement on Sume at $0.20.
- Captions unreadable on busy footage: dim the clip, then burn
Busy B-roll can swallow burned captions. Sume can dim the clip with video-filter, then burn captions with colour overrides, for about $0.22 a clip.
- Check character drift across AI shots with video_frames
Pull up to 24 evenly spaced stills from each AI clip with the unbilled video_frames route and compare faces and products before you stitch shots together.
- Check an AI take says your script: hypit align and unmatched words
hypit align pairs each script token with the transcript words of a generated take, and lists unmatched words. What it measures and what it does not do.
Written by Sume