Caption a Reel after First Draft: burn once on the final cut

First Draft shifts where every word lands, so burn captions only on the final cut and restyle with source_caption_id. The $0.20 job, errors and corrections.

4 min readSume
All posts

Burn captions last. First Draft removes pauses and trims clips, so every word after the first cut lands earlier than it did in the raw take, and captions timed on raw footage will drift. Run Sume video captions once on the finished cut, then restyle from the saved caption id instead of transcribing again.

TechCrunch describes First Draft as cutting pauses and trimming the clips you selected (TechCrunch, read 2026-10-07). The practical consequence is order of operations, not a feature.

The order that avoids re-doing work

  • 1. Cut: First Draft, or your own trims.
  • 2. Review the cut by eye. Move or delete anything the tool got wrong.
  • 3. Export the final MP4 and make it a public HTTPS URL that Sume can fetch (video captions take a public video_url).
  • 4. POST /v1/video-captions once, with style and optional design overrides.
  • 5. Save the returned caption id. If you want another look, send source_caption_id with a new style.

Why restyling is the cheap path

The docs say a restyle reuses the source video of the first caption together with the word timings Sume already has, so speech-to-text does not run a second time. The price does not change, because a restyle is still a render: each accepted standalone caption job is $0.20 for a video of up to 60 seconds, under the current fixed estimate.

That makes a look test cheap in time but not free in money. Three looks on one Reel is three renders.

Caption jobs on one 45-second Reel at the Sume public rate (read 2026-10-07)
ActionRequest body keySpeech-to-text again?Price
First burnvideo_urlYes, or script_text alignment$0.20
Second looksource_caption_id + styleNo$0.20
Third looksource_caption_id + styleNo$0.20
Three looks total$0.60

When the words are wrong

If a name or a product term comes out wrong, send words with the corrected text on the restyle, or send script_text on the first burn so Sume aligns your script to the speech-to-text timings. Send only one of script_text, words, cues and segments.

A silent clip has nothing to transcribe and fails with caption_no_speech; for a silent clip send cues with text, start and end in seconds.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: final-cut-look-2" \
  -d '{ "source_caption_id": "vc_123", "style": "punch" }'

What goes wrong if you caption first

Suppose you caption the raw 50-second clip, then a first-pass tool removes four pauses of about half a second each. Every word after the first cut is now about two seconds early or late against the burned text, and a burned caption cannot be moved; you have to render again. That second render is another $0.20, and a speech-to-text run too if you start from the new file rather than from a saved caption id.

The failure is quiet. A caption off by a second looks fine in a thumbnail and bad in motion, which is why it survives review. Burn once, after the cut is final, and look at the result with the sound on.

A habit that keeps the order

Name the files by stage: raw, cut, final, final-captioned. A caption job only ever takes a final. If a change is needed after captions, make it in the cut, re-export, and burn again; do not try to patch the captioned file. The one exception is a text correction, where source_caption_id plus words fixes the spelling without a second transcription.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume