Accessible transcript for a video page: Sume transcript, then review

Get a timed transcript of your video from Sume's video inspect for $0.01 a minute, then correct it by ear. Silent clips fail, and no output is pre-certified.

5 min readSume
All posts

To get a text transcript for a video page, call POST /v1/video-inspect with transcribe: true on a clip already in your workspace's media. The transcript comes back with text, word timings and, if you ask, sentence segments, at $0.01 per audio minute. Treat it as a first draft: the W3C notes that automatic captions often are wrong and need human review.

What does the transcript contain?

The inspect resource carries probe, frames, a transcript when requested (text, words[], optional sentence segments[] and an audio_url) and warnings[]. Probe and stills are unbilled; only the transcript half reserves. Omitting duration_seconds reserves one minute, and the hint is capped at 600 seconds.

Set language_code (for example en or ko) as a speech-to-text hint, or leave it out for auto-detection. Adding segmentation.mode: "sentence" returns gapless sentence segments, and silence_split_seconds between 0.2 and 3 tunes where lines break.

What does the request look like?

The source must be this workspace's media.sume.com clip, no longer than 1800 seconds. Import first with POST /v1/media-imports. Default mode is sync: the handler waits up to 30 seconds, then answers 200 with the finished inspect or 202 with a job to poll.

curl -X POST https://api.sume.com/v1/video-inspect \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: transcript-lesson-3" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/lesson3.mp4",
    "frames": false,
    "transcribe": true,
    "language_code": "en",
    "segmentation": { "mode": "sentence" }
  }'

What does the W3C expect from a good caption or transcript?

The W3C WAI captions page describes captions as speech plus the non-speech audio needed to understand the content, with speaker identification and synchronisation, and says automatic captions need to be edited. A transcript from Sume covers the speech only.

Quality gap checklist for the draft (read 2026-10-02)
RequirementFrom Sume transcriptYour review step
Accurate wordsDraft onlyListen and correct names and numbers
Non-speech soundsNot includedAdd bracketed notes where they matter
Speaker identityNot includedAdd names where unclear
TimingWord timings and sentence segmentsSpot-check against the audio
Line breaksSentence segmentsAdjust by hand

What fails, and how do you avoid it?

A clip with no audio track fails inspect_source_has_no_audio; check probe.has_audio first, since an inspect with frames: false is enough. The fields language_code, segmentation and duration_seconds without transcribe: true return 400 video_inspect_transcribe_required. Off-host URLs are rejected at admit.

If a recording is long, detach the audio once with Audio detach, which can output 16 kHz mono wav for speech-to-text at $0.01 per job.

What do you do with the transcript?

Publish it on the video page for people who read rather than watch, and reuse the segments as caption cues if you also want open captions burned in. Sume does not publish the page or host a caption track; that stays on your site. Keep the reviewed version, not the raw draft.

How do you keep the review manageable?

Review in the same order as the clip. Correct names, numbers and technical terms first, since those are where speech-to-text slips most, then fix line breaks. Word timings let you jump to a questionable word: each entry in words[] carries timing, so you can seek straight to the spot.

For a series, keep a short glossary of names and terms next to the transcripts. If you also burn captions, pass your corrected script as script_text so the burned wording matches your reviewed text and not the raw recognition.

When should you caption as well?

A transcript helps readers and search, but it does not replace captions for someone watching the video. If you want both, build the transcript first, correct it, and then use the corrected sentence segments as cues for a caption burn with Video captions, at $0.20 per clip up to 60 seconds under the current estimate. That way one review pass serves both outputs.

What does this not guarantee?

Nothing here certifies a video as accessible under any law or guideline. It saves the typing; the listening and correcting is the part that makes a transcript trustworthy.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume