Test three caption looks on one video without re-transcribing

Burn a video once, then restyle it twice with source_caption_id. Word timings are reused, so no second speech-to-text. Each render is still $0.20.

5 min readSume
All posts

To try three caption looks on one video, burn it once with POST /v1/video-captions, then send two more requests with source_caption_id set to the first job's caption id and a different style each time. Sume reuses the source video and the word timings it already has, so speech-to-text does not run again. Each restyle is still a render and still billed at the fixed $0.20 for clips up to 60 seconds.

So you save time and keep the wording identical across the variants, but not money. Three looks cost three renders.

What does the restyle request send?

Instead of video_url, pass source_caption_id with the id of the earlier caption resource. Add style, and optionally design overrides. Pass words only if you want to correct the wording; it is not needed for a pure restyle.

The source_caption_id value is the resource id from the first job's response. The docs write it as a placeholder, so copy it from your own result rather than building it.

import asyncio, json, os, urllib.request

BASE = "https://api.sume.com/v1/video-captions"

def post(body, idem):
    req = urllib.request.Request(BASE, data=json.dumps(body).encode(), method="POST",
        headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
                 "Content-Type": "application/json", "Idempotency-Key": idem})
    return json.load(urllib.request.urlopen(req))

async def main(caption_id):
    for i, style in enumerate(["punch", "tiktok-green"]):
        out = await asyncio.to_thread(post,
            {"source_caption_id": caption_id, "style": style}, f"look-test-{i}")
        print(style, out.get("request_id"))

if not os.environ.get("SUME_API_KEY"):
    raise SystemExit("set SUME_API_KEY")
asyncio.run(main(os.environ["FIRST_CAPTION_ID"]))

Which looks are worth comparing?

Omit style and the wording decides: slam for Latin text, black-outline for Korean. Named looks include punch and tiktok-green for Latin speech, and the Hangul set for Korean. design is not supported on punch or tiktok-green, so tune colour and placement on slam or a Hangul style.

Restyle options, read 2026-10-02
OptionWorks withNote
styleAll named stylesLatin copy cannot use a Hangul style font
design.colors.activeslam and Hangul stylesSets the spoken-word tint
design.placement.anchor_ratioslam and Hangul stylesLine centre as a fraction of frame height
fontHangul styles onlyA Latin style with a font returns 400

What does a restyle not reuse?

It reuses the transcription, not the earlier burn. Each variant is rendered fresh from the source video, so the first caption's look does not carry over. Wait for the first job to finish before you submit variants, since the restyle points at its result.

A restyle does not re-run recognition. If the transcript is wrong, the docs say to pass corrected words alongside source_caption_id; otherwise start again from the video.

How do I judge the winner?

Sume does not measure which look performs better. Post the variants where you can see retention, and judge them there. Before you do, read all three on a phone at normal size, because a look that is clear on a desktop preview can be hard to read at arm's length.

Poll each job and read the results as described in Jobs and results. Source videos must be fetchable public HTTPS URLs; see Media inputs for importing yours.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume