Get transcript text from a captioned video: caption jobs return none

A Sume caption job returns the burned video, not the transcript. For text and word times run STT or video inspect with transcribe. Prices and a recipe.

5 min readSume
All posts

A Sume video caption job does not hand back the transcript. The caption docs say the resource returns the public-safe status, the style and the captioned video_url and artifacts, and that raw transcripts, renderer internals and signed source URLs are not part of the public contract. If you need the text, ask for it separately: run Sume STT on the audio, or call video inspect with transcribe: true on the clip.

This matters for anyone who wants subtitles and a blog post, search index or translation from the same video.

Two routes to the text

Route A is video inspect. POST /v1/video-inspect with transcribe: true runs Sume STT 1.0 on a clip you already host on media.sume.com, and the result includes transcript with text, words[], optional sentence segments[] and an audio_url. The transcript line is $0.01 per audio minute. Without duration_seconds Sume reserves one minute, and the hint maxes at 600 seconds. A clip with no audio fails with inspect_source_has_no_audio.

Route B is STT directly. Detach the audio with POST /v1/audio-detach, then send the WAV to POST /v1/stt-1.0/transcribe. That costs $0.01 for the detach plus $0.01 per audio minute, which is a cent more than route A for a one minute clip but gives you the audio file as a separate artifact.

Getting text from one 45 second clip, rates from Sume docs read 2026-10-07
RouteCallsCost excluding compute reservationYou get
Caption job only1$0.20Burned video, no raw transcript
Video inspect with transcribe1$0.01 transcript linetext, words, optional segments, audio_url
Detach then STT2$0.01 + $0.01WAV artifact, text, words
Both caption and inspect2$0.21 plus inspect computeVideo and text

Keep the wording consistent

The cleanest pattern is to transcribe once, edit the text, and caption from your edited words. Run inspect or STT, fix names and numbers in the returned words, then send those as words or script_text on the caption job. That way the text in your blog post and the text on screen are the same text. If you caption first and transcribe second, the two runs may not match word for word, because they are separate recognizer calls.

Example

Inspect for text only, skipping stills with frames: false. Use the segmentation block if you want sentence rows.

curl -X POST https://api.sume.com/v1/video-inspect \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: inspect-text-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
    "frames": false,
    "transcribe": true,
    "duration_seconds": 60,
    "segmentation": { "mode": "sentence" }
  }'

What to store

Save the job id from each call and the text you approved. If a caption looks wrong later, you can re-burn with source_caption_id and a corrected words list instead of paying for recognition again. The transcript itself is yours to keep once returned; the caption resource will not return it later.

When you only need text

If you never plan to burn captions, do not pay for the caption job at all. A probe-only inspect is free, and adding transcribe: true adds only the per-minute transcript rate on top of the inspect's compute. For a 5 minute talk with duration_seconds set to 300, the transcript line is 5 x $0.01 = $0.05. Setting the hint matters: with it absent, Sume reserves only one minute for the transcript.

If the video is not hosted on Sume yet, import it first with POST /v1/media-imports; inspect and detach do not fetch from the open internet.

One last point on terminology. People say transcript, caption file and subtitles as if they were one thing. On Sume they are three outputs: a transcript is text with timings, a caption job is a rendered video, and a sidecar file such as SRT is something you build from the transcript yourself. Burned-in captions cannot be switched off by the viewer, so keep the transcript if you also want a toggle on another platform.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume