Accessible transcript for a video page: Sume transcript, then review
Get a timed transcript of your video from Sume's video inspect for $0.01 a minute, then correct it by ear. Silent clips fail, and no output is pre-certified.

To get a text transcript for a video page, call POST /v1/video-inspect with transcribe: true on a clip already in your workspace's media. The transcript comes back with text, word timings and, if you ask, sentence segments, at $0.01 per audio minute. Treat it as a first draft: the W3C notes that automatic captions often are wrong and need human review.
What does the transcript contain?
The inspect resource carries probe, frames, a transcript when requested (text, words[], optional sentence segments[] and an audio_url) and warnings[]. Probe and stills are unbilled; only the transcript half reserves. Omitting duration_seconds reserves one minute, and the hint is capped at 600 seconds.
Set language_code (for example en or ko) as a speech-to-text hint, or leave it out for auto-detection. Adding segmentation.mode: "sentence" returns gapless sentence segments, and silence_split_seconds between 0.2 and 3 tunes where lines break.
What does the request look like?
The source must be this workspace's media.sume.com clip, no longer than 1800 seconds. Import first with POST /v1/media-imports. Default mode is sync: the handler waits up to 30 seconds, then answers 200 with the finished inspect or 202 with a job to poll.
curl -X POST https://api.sume.com/v1/video-inspect \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: transcript-lesson-3" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/lesson3.mp4",
"frames": false,
"transcribe": true,
"language_code": "en",
"segmentation": { "mode": "sentence" }
}'What does the W3C expect from a good caption or transcript?
The W3C WAI captions page describes captions as speech plus the non-speech audio needed to understand the content, with speaker identification and synchronisation, and says automatic captions need to be edited. A transcript from Sume covers the speech only.
| Requirement | From Sume transcript | Your review step |
|---|---|---|
| Accurate words | Draft only | Listen and correct names and numbers |
| Non-speech sounds | Not included | Add bracketed notes where they matter |
| Speaker identity | Not included | Add names where unclear |
| Timing | Word timings and sentence segments | Spot-check against the audio |
| Line breaks | Sentence segments | Adjust by hand |
What fails, and how do you avoid it?
A clip with no audio track fails inspect_source_has_no_audio; check probe.has_audio first, since an inspect with frames: false is enough. The fields language_code, segmentation and duration_seconds without transcribe: true return 400 video_inspect_transcribe_required. Off-host URLs are rejected at admit.
If a recording is long, detach the audio once with Audio detach, which can output 16 kHz mono wav for speech-to-text at $0.01 per job.
What do you do with the transcript?
Publish it on the video page for people who read rather than watch, and reuse the segments as caption cues if you also want open captions burned in. Sume does not publish the page or host a caption track; that stays on your site. Keep the reviewed version, not the raw draft.
How do you keep the review manageable?
Review in the same order as the clip. Correct names, numbers and technical terms first, since those are where speech-to-text slips most, then fix line breaks. Word timings let you jump to a questionable word: each entry in words[] carries timing, so you can seek straight to the spot.
For a series, keep a short glossary of names and terms next to the transcripts. If you also burn captions, pass your corrected script as script_text so the burned wording matches your reviewed text and not the raw recognition.
When should you caption as well?
A transcript helps readers and search, but it does not replace captions for someone watching the video. If you want both, build the transcript first, correct it, and then use the corrected sentence segments as cues for a caption burn with Video captions, at $0.20 per clip up to 60 seconds under the current estimate. That way one review pass serves both outputs.
What does this not guarantee?
Nothing here certifies a video as accessible under any law or guideline. It saves the typing; the listening and correcting is the part that makes a transcript trustworthy.
Sources
Related posts
More in Use cases
- AI image with a working QR code: Muse Image uses code, others guess
Meta says Muse Image runs code to make accurate QR codes. Image models just draw one that may not scan. Generate the art, then add a real QR yourself.
- AI thumbnail variants in one call: n=4, pick one, then polish
Ask for four thumbnail options in one POST /v1/images call with n up to 4, pick the best by eye, then run a polish edit. Cost per step and when n is a waste.
- AI-written script on a public-interest topic: Article 50 text label
Article 50(4) text labels cover published public-interest text with no human review or editorial control. What counts as review, per the Commission.
- Amazon DVA+ rolls out from late October: get your video assets ready
Amazon says DVA+ begins rolling out in late October and advertisers approve creatives before launch. A checklist for preparing holiday video assets with Sume.
Written by Sume