Caption a Recast video: burn captions after the person swap
Run h3-max-recast first, then send its output to POST /v1/video-captions. The kept source audio gives the captions speech to transcribe.

To caption a person-swapped video on Sume, finish the Recast job first and then submit its output URL to POST /v1/video-captions. Order matters: Recast re-renders the footage, so captions burned before the swap would be re-rendered with it, while captions burned after sit on the final pixels. Recast keeps the source audio, so the output has the same speech the captions need to transcribe (Video Router docs and Video captions docs, read 2026-10-03).
The two calls are independent jobs. You submit Recast, wait for a terminal state, take the Sume-hosted video URL from the result, and submit that URL to captions.
Why captions go last
A Recast clip is a new render at 768p or 1080p. Anything drawn on the source, including captions, is part of the source frame and gets treated as scene content. That can leave text distorted or duplicated. Burning captions on the output avoids it entirely and also lets you pick the style once, after you know which cast you are shipping. If you produce five variants of one ad with five different people, you will run captions five times, but the style, language and font stay the same across all of them.
There is a second reason. The captions endpoint transcribes with speech-to-text unless you pass your own text. The docs say a clip with no audible speech fails as caption_no_speech with next_action: use_overlay_captions, so a silent output is a dead end for automatic captions. Because Recast keeps the source audio, a source with speech gives you an output with speech.
The two requests
The first call recasts. The second burns captions on the finished clip. The caption body needs only video_url; style, font, language, script_text and cues are optional. Omitting style lets the wording decide: slam for Latin text, black-outline for Korean.
# 1) person swap
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" -H "Content-Type: application/json" \
-H "Idempotency-Key: recast-cap-001" \
-d '{"model":"h3-max-recast","video_url":"https://example.com/ad.mp4","reference_image_urls":["https://example.com/host.jpg"],"resolution":"1080p"}'
# 2) captions on the finished clip (use the media.sume.com URL from the result)
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" -H "Content-Type: application/json" \
-d '{"video_url":"https://media.sume.com/artifacts/artf_demo/recast.mp4","style":"slam","language":"en"}'Choices that affect the result
Use the language field as a hint to speech-to-text rather than as a style switch; the docs say it never selects the style or the font. If the source script is known, pass it as script_text so the captions align to your wording instead of a transcription. If you want captions that were authored by a person, pass cues with text, start and end, which skips transcription.
| Situation | What to send | Why |
|---|---|---|
| Source had clear speech | video_url only | Speech-to-text transcribes the kept audio |
| You have the script | video_url + script_text | Aligns your wording to the speech |
| Source had no speech | video_url + cues | Silent clips fail as caption_no_speech |
| Brand colors needed | design overrides | Merge over the style's own tokens |
| Korean copy | Omit style or name korean-ad | Latin default would burn tofu |
Edge cases worth testing once
Run one clip through the full chain before you batch. Three things tend to surprise people. First, the text in the captions comes from speech-to-text on the Recast output, so a noisy source produces noisy captions; if accuracy matters, pass script_text. Second, the caption job is its own reservation, so a failed Recast never costs you a caption and a failed caption never costs you a Recast; the two bills are independent. Third, the caption endpoint wants a public HTTPS video_url, so use the Sume-hosted URL that the Recast result returns rather than a signed or private link, which the docs say are rejected.
If your platform shows its own auto-captions, you may not need to burn any in. Burned captions are part of the picture and cannot be turned off by the viewer, which suits short vertical video and is a poor choice where accessibility settings or translation tracks matter. Decide per channel, not per project.
A short publishing checklist
Before you post a captioned swap, review it. Check that the replaced person is consistent across every cut, that the words in the captions match what is said and that the captions clear the platform's safe area for your format. Add whatever disclosure your platform requires for AI-altered footage; Sume documents the tools and does not decide platform policy for you.
If you need to cut the clip down for a feed that wants a shorter runtime, do it with video trim before captions, not after, so the captions are timed to the final cut.
- Recast first, captions second, trim before both if the source is long.
- Poll or use a webhook between steps; each is a separate job.
- Keep one
Idempotency-Keyper step.
Sources
Related posts
More in Media tools
- Caption micro-drama dialogue per shot: $0.20 per clip up to 60 s
Video captions cost $0.20 per job for clips up to 60 seconds. Caption each shot before assembly; a silent clip fails with caption_no_speech. Budget math.
- Caption line breaks: Netflix bottom-heavy pyramid rule vs Sume cards
Netflix wants two-line captions broken as a bottom-heavy pyramid. Sume caption cards break by phrasing settings. How to approximate the rule and where it ends.
- Captions for applause and music cues: author the cue text yourself
ElevenLabs Scribe can tag laughter and applause. On Sume, a silent clip returns caption_no_speech, so supply your own cues.
- Check a product-swap video edit for the old product with frames
After a prompted product swap on a video, pull matching stills from the source and the edit and compare them to catch the old product. Sume docs for each step.
Written by Sume