Caption a Reel after First Draft: burn once on the final cut
First Draft shifts where every word lands, so burn captions only on the final cut and restyle with source_caption_id. The $0.20 job, errors and corrections.

Burn captions last. First Draft removes pauses and trims clips, so every word after the first cut lands earlier than it did in the raw take, and captions timed on raw footage will drift. Run Sume video captions once on the finished cut, then restyle from the saved caption id instead of transcribing again.
TechCrunch describes First Draft as cutting pauses and trimming the clips you selected (TechCrunch, read 2026-10-07). The practical consequence is order of operations, not a feature.
The order that avoids re-doing work
- 1. Cut: First Draft, or your own trims.
- 2. Review the cut by eye. Move or delete anything the tool got wrong.
- 3. Export the final MP4 and make it a public HTTPS URL that Sume can fetch (video captions take a public
video_url). - 4.
POST /v1/video-captionsonce, withstyleand optionaldesignoverrides. - 5. Save the returned caption id. If you want another look, send
source_caption_idwith a newstyle.
Why restyling is the cheap path
The docs say a restyle reuses the source video of the first caption together with the word timings Sume already has, so speech-to-text does not run a second time. The price does not change, because a restyle is still a render: each accepted standalone caption job is $0.20 for a video of up to 60 seconds, under the current fixed estimate.
That makes a look test cheap in time but not free in money. Three looks on one Reel is three renders.
| Action | Request body key | Speech-to-text again? | Price |
|---|---|---|---|
| First burn | video_url | Yes, or script_text alignment | $0.20 |
| Second look | source_caption_id + style | No | $0.20 |
| Third look | source_caption_id + style | No | $0.20 |
| Three looks total | $0.60 |
When the words are wrong
If a name or a product term comes out wrong, send words with the corrected text on the restyle, or send script_text on the first burn so Sume aligns your script to the speech-to-text timings. Send only one of script_text, words, cues and segments.
A silent clip has nothing to transcribe and fails with caption_no_speech; for a silent clip send cues with text, start and end in seconds.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: final-cut-look-2" \
-d '{ "source_caption_id": "vc_123", "style": "punch" }'What goes wrong if you caption first
Suppose you caption the raw 50-second clip, then a first-pass tool removes four pauses of about half a second each. Every word after the first cut is now about two seconds early or late against the burned text, and a burned caption cannot be moved; you have to render again. That second render is another $0.20, and a speech-to-text run too if you start from the new file rather than from a saved caption id.
The failure is quiet. A caption off by a second looks fine in a thumbnail and bad in motion, which is why it survives review. Burn once, after the cut is final, and look at the result with the sound on.
A habit that keeps the order
Name the files by stage: raw, cut, final, final-captioned. A caption job only ever takes a final. If a change is needed after captions, make it in the cut, re-export, and burn again; do not try to patch the captioned file. The one exception is a text correction, where source_caption_id plus words fixes the spelling without a second transcription.
Sources
Related posts
More in Media tools
- Caption a talking-head video: style, placement and phrasing
Caption a talking-head clip with the Sume API: pick a style, move the line off the face, set words per card. One request, $0.20 for up to 60 seconds.
- How do I caption a Shorts series in two languages in one batch?
Run one Sume caption job per episode and language: $0.20 each for clips up to 60 seconds, so a 10-episode season in two languages is $4.00.
- Captions for a video with loud background music: script text or cues
Music can bury speech and trip speech-to-text. Sume documents three ways to supply your own wording: script_text, a words array or cues. Which to pick.
- ChatGPT Sora clip as reference video: Omni takes 3 s each, trim
Sora 2 left the API but a ChatGPT clip can still seed a new render. Gemini Omni Flash 1.1 on Sume takes up to 3 reference videos, each 3 s at most.
Written by Sume