Restyle ad captions without transcribing the video twice

Pass source_caption_id to Sume's video-captions endpoint to re-burn an ad in a new style, reusing the first run's word timings with no second speech-to-text.

5 min readSume
All posts

To try a different caption style on the same ad, call POST /v1/video-captions with source_caption_id instead of video_url, plus the new style. Sume reuses the earlier caption's source video and its word timings, so no second speech-to-text runs. A restyle is still a render, so billing is unchanged per the Video captions docs.

Request shape

The doc's example body is { "source_caption_id": "...", "style": "black-outline" }. Pass words alongside it only to correct wording.

Restyle fields, from the Sume docs (read 2026-10-03).
FieldRole in a restyle
source_caption_idReplaces video_url; reuses source video and timings
styleThe new look
wordsOptional wording corrections

Why it suits A/B work

Caption look is a clean single variable. Burn one transcript in two styles and compare, without the transcription varying between arms.

Cost and scope

The standalone caption job lists at $0.20 per job for clips up to 60 seconds (list prices from Sume's catalog rate card (GET /v1/catalog), confirmed in the repo's catalog file read 2026-10-03). Inline captions on an avatar video are an add-on to that video's estimate, not a separate caption job.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume