Captions with sound cues for deaf viewers: W3C checklist on Sume

W3C says captions carry speech and non-speech sound. Sume's STT burn covers speech only, so here is how to author the full cue list and burn it.

5 min readSume
All posts

Burned-in captions from Sume cover what is said, not what is heard. The W3C describes captions as a text version of "the speech and non-speech audio information needed to understand the content", so a creator who wants accessible captions has to add the sound cues themselves and burn the complete list as authored cues with Video captions.

This post walks through that workflow: transcribe once, merge in sound cues, burn, and review. It does not claim that Sume output meets any legal standard; that judgement stays with you.

What does the W3C say a caption must contain?

The W3C WAI captions page is the reference for what follows. These are the parts that matter for a small team:

  • Speech and non-speech audio: the W3C definition covers both. A door slam or a phone ring is our own example of non-speech audio.
  • Speaker identification: the page shows speaker labels in WebVTT, so label speakers where the viewer cannot tell who is talking.
  • Synchronisation: captions are timed to the audio.
  • Accuracy: the page says automatically generated captions do not meet accessibility requirements unless confirmed fully accurate, and usually need significant editing.

What does Sume's speech-to-captions path give you?

When you send only video_url (or script_text), Sume runs speech-to-text and burns the words it hears. That path needs audible speech: a silent clip fails as caption_no_speech. It produces no sound cues and no speaker names, because the transcript is words and timings only.

The other path is authored overlay copy. You pass cues (or segments), each with text, start and end in seconds, and Sume burns exactly that text at those times with no speech-to-text. The fields script_text, words, cues and segments are mutually exclusive, so a cue list has to hold everything you want on screen, speech included.

How do you build a complete cue list?

Get speech timings from Video inspect. It needs a media.sume.com artifact or asset of your workspace (import other clips first with POST /v1/media-imports) and an Idempotency-Key. Set transcribe: true with segmentation.mode: "sentence" and it returns gapless sentence segments[], shaped like caption lines. Transcription is billed at $0.01 per audio minute; probe and stills are not. Then:

1. Copy each sentence segment into a cue with its start and end.

2. Insert sound cues in square brackets, such as [door slams], in the gaps where the sound happens.

3. Prefix a speaker name when the picture does not make it obvious, such as MAYA:.

4. Correct the words by ear. Sume does not check your edits against the audio.

The docs describe no separate treatment for bracketed text, so sound cues burn in the same style as speech. Pick one style that stays legible for both.

Cue sources and what Sume burns (read 2026-10-02)
ElementWhere it comes fromBurned by Sume?
Spoken wordsVideo inspect sentence segments, edited by youYes, as cues
Sound cuesYou write themYes, as cues
Speaker namesYou write themYes, as cue text
Sidecar caption fileNot producedNo: SRT uploads are unsupported

What does the burn call look like?

The clip must be a fetchable public HTTPS video URL. Each accepted standalone job is $0.20 for videos up to 60 seconds under the current fixed estimate; check live pricing in GET /v1/catalog.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: access-captions-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/example/clean.mp4",
    "style": "black-outline",
    "cues": [
      {"text": "[door slams]", "start": 0.0, "end": 1.2},
      {"text": "MAYA: Welcome back to the studio.", "start": 1.4, "end": 3.6}
    ]
  }'

How should you review the result?

Play the burned clip with the sound on and then with the sound off. With sound on, check that each word matches what you hear and that each bracketed cue appears when the sound does. With sound off, check that the story still makes sense: if a viewer would miss a plot point, a cue is missing.

Keep cues short. A bracketed note such as [phone rings] reads in a glance, while a sentence of description pulls the eye away from the picture. If a cue needs fixing after the burn, the speech timings stay yours: change the cue list and run the caption job again, or restyle with source_caption_id when you only want a new look and not new words.

What does the whole pipeline cost?

For a 45-second clip the published rates are $0.01 for one minute of transcript through Video inspect and $0.20 for the burn, so about $0.21 per pass. Review time is the real cost, and it is the part that cannot be automated away. Rates can change; confirm them in GET /v1/catalog before a batch.

What does Sume not do here?

Sume burns captions into the picture, so viewers cannot turn them off, and it does not produce a caption file for a platform's own caption track. It does not detect non-speech sounds or identify speakers for you. If a platform needs a sidecar track, that is a separate tool in your pipeline; see burned-in versus sidecar captions for the trade-off.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume