Restyle avatar clip captions with source_caption_id, no re-transcribe
To try a second caption look on an avatar video, send source_caption_id instead of the video URL. Sume reuses the word timings. Cost, errors and a worked flow.
To test another caption style on the same avatar clip, call the standalone captions endpoint with source_caption_id and a new style, not video_url. The docs say Sume reuses the source video of that caption job with the word timings it already has, so speech-to-text does not run again. The price does not change, because a restyle is still a render. Each accepted standalone caption job is a flat 0.20 dollars for videos up to 60 seconds.
The flow
- Render the avatar clip with captions off so you have a clean public MP4.
- Create a caption job with
video_urland a first style. Keep the caption id from the response. - To try a second look, send
{ "source_caption_id": "...", "style": "punch" }. - If the transcript has a wrong word, send
wordswith it to correct the text only. - Compare the results in the destination app, then keep one.
What it costs to try three looks
Three caption renders cost three flat charges. Compare that with re-rendering the avatar clip, which is billed per second, and the point is clear: caption experiments are cheap, avatar re-rolls are not.
| Step | Cost |
|---|---|
| Avatar clip, 30 s at plus, no product image | $7.35 |
| First caption render | $0.20 |
| Restyle 1 via source_caption_id | $0.20 |
| Restyle 2 via source_caption_id | $0.20 |
| Total | $7.95 |
Pitfalls
Do not send video_url and source_caption_id together; the docs describe them as alternatives. If you send script_text for alignment and the clip's speech does not match, the job can fail with script_alignment_mismatch or script_alignment_failed, and the docs recommend simplifying the text or omitting it. Only one of script_text, words, cues and segments is accepted per request.
punch and tiktok-green do not support design overrides, so a restyle to those looks cannot be tuned further. Korean text on slam, punch or tiktok-green returns 400 caption_hangul_text_latin_style; this is a design guard, not a bug.
When inline captions are enough
If you need one standard look and never restyle, inline captions on the avatar request are simpler: one call, no second job. Their cost sits in the avatar estimate, they work to 60 seconds, and a failure in the caption stage is soft, leaving a clean video with captions.status=failed. Pick the standalone route when you want control, and the inline one when you want fewer steps.
Keeping the record straight
Store the first caption id next to the avatar-video id, and store each restyle's id too. When a stakeholder asks for 'the one with the yellow word', you can find it without rendering again. Use the style name in your own label, because the raw ids are opaque. Archive the rejected looks, since they cost money to make.
Sources
Related posts
More in Developers
- Retry a failed Format create with the same Idempotency-Key
After a 402 or 503 on a Format create, Sume releases the Idempotency-Key. Fix the cause, then retry with the same key instead of minting a new one.
- Rotate the Sume webhook signing secret without dropping a delivery
Upgrade the verifier first, rotate with POST /v1/webhooks/signing-secret/rotate, deploy the new secret inside the 24-hour two-signature window, then confirm it.
- Route Sume run webhooks by event: format, action and agent terminal
Run webhooks use one terminal event per family: action.run.terminal, format.run.terminal, agent.run.terminal. Route on event, then branch on outcome.
- Ruby Net::HTTP: create a Sume bulk queue and poll to an exit code
A 30-line Ruby script with only the standard library: create a Sume bulk queue from items.json, back off the poll, skip 429 and 503, and exit 1 on failed items.
Written by Sume