Caption a Recast video: burn captions after the person swap

Run h3-max-recast first, then send its output to POST /v1/video-captions. The kept source audio gives the captions speech to transcribe.

6 min readSume
All posts

To caption a person-swapped video on Sume, finish the Recast job first and then submit its output URL to POST /v1/video-captions. Order matters: Recast re-renders the footage, so captions burned before the swap would be re-rendered with it, while captions burned after sit on the final pixels. Recast keeps the source audio, so the output has the same speech the captions need to transcribe (Video Router docs and Video captions docs, read 2026-10-03).

The two calls are independent jobs. You submit Recast, wait for a terminal state, take the Sume-hosted video URL from the result, and submit that URL to captions.

Why captions go last

A Recast clip is a new render at 768p or 1080p. Anything drawn on the source, including captions, is part of the source frame and gets treated as scene content. That can leave text distorted or duplicated. Burning captions on the output avoids it entirely and also lets you pick the style once, after you know which cast you are shipping. If you produce five variants of one ad with five different people, you will run captions five times, but the style, language and font stay the same across all of them.

There is a second reason. The captions endpoint transcribes with speech-to-text unless you pass your own text. The docs say a clip with no audible speech fails as caption_no_speech with next_action: use_overlay_captions, so a silent output is a dead end for automatic captions. Because Recast keeps the source audio, a source with speech gives you an output with speech.

The two requests

The first call recasts. The second burns captions on the finished clip. The caption body needs only video_url; style, font, language, script_text and cues are optional. Omitting style lets the wording decide: slam for Latin text, black-outline for Korean.

# 1) person swap
curl -X POST https://api.sume.com/v1/video-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" -H "Content-Type: application/json" \
  -H "Idempotency-Key: recast-cap-001" \
  -d '{"model":"h3-max-recast","video_url":"https://example.com/ad.mp4","reference_image_urls":["https://example.com/host.jpg"],"resolution":"1080p"}'

# 2) captions on the finished clip (use the media.sume.com URL from the result)
curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" -H "Content-Type: application/json" \
  -d '{"video_url":"https://media.sume.com/artifacts/artf_demo/recast.mp4","style":"slam","language":"en"}'

Choices that affect the result

Use the language field as a hint to speech-to-text rather than as a style switch; the docs say it never selects the style or the font. If the source script is known, pass it as script_text so the captions align to your wording instead of a transcription. If you want captions that were authored by a person, pass cues with text, start and end, which skips transcription.

Caption inputs after a Recast job, read 2026-10-03
SituationWhat to sendWhy
Source had clear speechvideo_url onlySpeech-to-text transcribes the kept audio
You have the scriptvideo_url + script_textAligns your wording to the speech
Source had no speechvideo_url + cuesSilent clips fail as caption_no_speech
Brand colors neededdesign overridesMerge over the style's own tokens
Korean copyOmit style or name korean-adLatin default would burn tofu

Edge cases worth testing once

Run one clip through the full chain before you batch. Three things tend to surprise people. First, the text in the captions comes from speech-to-text on the Recast output, so a noisy source produces noisy captions; if accuracy matters, pass script_text. Second, the caption job is its own reservation, so a failed Recast never costs you a caption and a failed caption never costs you a Recast; the two bills are independent. Third, the caption endpoint wants a public HTTPS video_url, so use the Sume-hosted URL that the Recast result returns rather than a signed or private link, which the docs say are rejected.

If your platform shows its own auto-captions, you may not need to burn any in. Burned captions are part of the picture and cannot be turned off by the viewer, which suits short vertical video and is a poor choice where accessibility settings or translation tracks matter. Decide per channel, not per project.

A short publishing checklist

Before you post a captioned swap, review it. Check that the replaced person is consistent across every cut, that the words in the captions match what is said and that the captions clear the platform's safe area for your format. Add whatever disclosure your platform requires for AI-altered footage; Sume documents the tools and does not decide platform policy for you.

If you need to cut the clip down for a feed that wants a shorter runtime, do it with video trim before captions, not after, so the captions are timed to the final cut.

  • Recast first, captions second, trim before both if the source is long.
  • Poll or use a webhook between steps; each is a separate job.
  • Keep one Idempotency-Key per step.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume