Test two caption looks on one Reel with source_caption_id
Restyling captions with source_caption_id reuses the first caption job's video and word timings, so no second transcription runs. Billing stays one render each.

To compare two caption looks on the same Reel, create the first caption job normally, then pass its id as source_caption_id with a different style instead of resending the video. Sume's video captions docs say this reuses that caption's source video and the word timings it already has, so no second speech-to-text runs. Billing is unchanged: a restyle is still a render. The saving is time and consistency, since both versions share one transcript and any correction you made to it.
Instagram's Edits app advertises a library of caption styles (read 2026-10-03). Here the styles are Sume's: slam, punch, tiktok-green and korean-ad, plus Hangul identities.
The request shapes
The first call takes video_url, a public HTTPS URL, and an optional style. The restyle takes source_caption_id and style, and you can pass words alongside it to correct wording. I did not find the exact field name that carries a caption id in the create response in the docs I read, so the script reads id and falls back to request_id; print the response once and confirm which one your restyle accepts.
import json, os, urllib.request
API = "https://api.sume.com/v1"
def post(path, body, key):
req = urllib.request.Request(
f"{API}{path}",
data=json.dumps(body).encode(),
headers={
"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Content-Type": "application/json",
"Idempotency-Key": key,
},
method="POST",
)
with urllib.request.urlopen(req) as res:
return json.load(res)
first = post(
"/video-captions",
{"video_url": os.environ["SUME_PUBLIC_CLIP_URL"], "style": "slam", "language": "en"},
"look-test-slam-v1",
)
caption_id = first.get("id") or first["request_id"]
print("first", caption_id)
second = post(
"/video-captions",
{"source_caption_id": caption_id, "style": "punch"},
"look-test-punch-v1",
)
print("second", second.get("id") or second["request_id"])
What to compare
For an English test, slam against punch is the cleanest pair: both are Latin-friendly, and differences in watch-through are about the look, not the words. Change only the style between the two versions; if you also change the font, the placement and the colour, you will not know which one moved the result.
Decide the metric before you post. Instagram shows you watch time and replays in Insights; pick one, publish both versions as separate Reels at comparable times, and compare after the same number of days. Two posts is a weak test, so read the result as a hint.
| Style | Notes (Sume docs) |
|---|---|
| slam | Default for Latin text |
| punch | design overrides not supported |
| tiktok-green | design overrides not supported |
| korean-ad | Hangul karaoke look; pair with language: ko |
Limits worth knowing
Captions bill $0.20 per job for clips up to 60 seconds, and a restyle counts as a render, so two looks cost two jobs. If the wording has errors, correct it once with words on the first or the restyle call, and the fix carries to later restyles built on that caption. script_text, words, cues and segments are mutually exclusive, so pick one correction method per request.
If an id is rejected, the usual cause is passing a job id where a caption id is expected, or a restyle against a caption that never completed. Read the job result and use the identifier it reports.
Keeping the experiment honest
A restyle test is easy to over-read. Both versions share the same footage and wording, which removes one source of noise, but audience, time of day and the order in which followers see them are still different. Post the two versions on different days of the same weekday if you can, and avoid reusing the same hook in the same week as another test.
Write down what you expect before you look at Insights. If you predicted that the heavier style would hold viewers longer and it did not, that is useful; a result you explain afterwards is mostly a story. When both versions land close together, keep the cheaper or faster one to produce, since the style then does not matter much for that audience.
Finally, save the caption id and style next to the post date in your notes. Months later, you will want to know which look you used, and the ids are the one thing Sume can use to reproduce it.
Sources
Related posts
More in Developers
- How to test a webhook URL before a Sume Format run uses it
POST /v1/webhooks/test-deliveries sends a signed webhook.test event to your URL. See the scope, the response fields, and the secret check, with no paid run.
- TikTok Display API video query: read is_aigc on 20 posts per call
TikTok's Query Videos endpoint returns is_aigc and counts for up to 20 video ids per call. Short Python audit, and where Sume fits.
- TikTok oEmbed: embed a posted clip, or host the MP4 yourself?
TikTok's oEmbed endpoint turns a video URL into an embed blockquote. When that beats hosting the file, and what a Sume-made MP4 changes about the choice.
- Timeline audio concat limits: 20 parts, 1,800 s, one channel layout
Timeline audio concat accepts 1 to 20 parts, outputs up to 1,800 seconds and needs one channel layout. Limits, refusal codes and a request that works.
Written by Sume