Get transcript text from a captioned video: caption jobs return none
A Sume caption job returns the burned video, not the transcript. For text and word times run STT or video inspect with transcribe. Prices and a recipe.

A Sume video caption job does not hand back the transcript. The caption docs say the resource returns the public-safe status, the style and the captioned video_url and artifacts, and that raw transcripts, renderer internals and signed source URLs are not part of the public contract. If you need the text, ask for it separately: run Sume STT on the audio, or call video inspect with transcribe: true on the clip.
This matters for anyone who wants subtitles and a blog post, search index or translation from the same video.
Two routes to the text
Route A is video inspect. POST /v1/video-inspect with transcribe: true runs Sume STT 1.0 on a clip you already host on media.sume.com, and the result includes transcript with text, words[], optional sentence segments[] and an audio_url. The transcript line is $0.01 per audio minute. Without duration_seconds Sume reserves one minute, and the hint maxes at 600 seconds. A clip with no audio fails with inspect_source_has_no_audio.
Route B is STT directly. Detach the audio with POST /v1/audio-detach, then send the WAV to POST /v1/stt-1.0/transcribe. That costs $0.01 for the detach plus $0.01 per audio minute, which is a cent more than route A for a one minute clip but gives you the audio file as a separate artifact.
| Route | Calls | Cost excluding compute reservation | You get |
|---|---|---|---|
| Caption job only | 1 | $0.20 | Burned video, no raw transcript |
| Video inspect with transcribe | 1 | $0.01 transcript line | text, words, optional segments, audio_url |
| Detach then STT | 2 | $0.01 + $0.01 | WAV artifact, text, words |
| Both caption and inspect | 2 | $0.21 plus inspect compute | Video and text |
Keep the wording consistent
The cleanest pattern is to transcribe once, edit the text, and caption from your edited words. Run inspect or STT, fix names and numbers in the returned words, then send those as words or script_text on the caption job. That way the text in your blog post and the text on screen are the same text. If you caption first and transcribe second, the two runs may not match word for word, because they are separate recognizer calls.
Example
Inspect for text only, skipping stills with frames: false. Use the segmentation block if you want sentence rows.
curl -X POST https://api.sume.com/v1/video-inspect \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: inspect-text-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"frames": false,
"transcribe": true,
"duration_seconds": 60,
"segmentation": { "mode": "sentence" }
}'What to store
Save the job id from each call and the text you approved. If a caption looks wrong later, you can re-burn with source_caption_id and a corrected words list instead of paying for recognition again. The transcript itself is yours to keep once returned; the caption resource will not return it later.
When you only need text
If you never plan to burn captions, do not pay for the caption job at all. A probe-only inspect is free, and adding transcribe: true adds only the per-minute transcript rate on top of the inspect's compute. For a 5 minute talk with duration_seconds set to 300, the transcript line is 5 x $0.01 = $0.05. Setting the hint matters: with it absent, Sume reserves only one minute for the transcript.
If the video is not hosted on Sume yet, import it first with POST /v1/media-imports; inspect and detach do not fetch from the open internet.
One last point on terminology. People say transcript, caption file and subtitles as if they were one thing. On Sume they are three outputs: a transcript is text with timings, a caption job is a rendered video, and a sidecar file such as SRT is something you build from the transcript yourself. Burned-in captions cannot be switched off by the viewer, so keep the transcript if you also want a toggle on another platform.
Sources
Related posts
More in Developers
- Handle every Sume API error with one switch on next_action
Sume errors share one envelope. Branch on next_action, retryable and retry_after_seconds, and your client handles new codes without a code change. JS sample.
- Hindi speech to text API: Sume STT with language_code hi
Transcribe Hindi audio with Sume STT: send language_code hi, check the reported language, and review code-mixed speech. $0.01 per audio minute.
- A 3-minute Timeline render: poll job status, don't sleep a fixed time
Timeline renders are async by default. Submit with an Idempotency-Key, poll /v1/jobs/:id/status until terminal, then read /result. Cost: 3 minutes is $0.30.
- Image-to-video not starting on my photo: frame_images vs references
Your photo is a reference, not a first frame, when it goes in input_references. Use frame_images with first_frame on Sume /v1/videos to pin the opening shot.
Written by Sume