Restyle avatar clip captions with source_caption_id, no re-transcribe

To try a second caption look on an avatar video, send source_caption_id instead of the video URL. Sume reuses the word timings. Cost, errors and a worked flow.

5 min readSume
All posts

To test another caption style on the same avatar clip, call the standalone captions endpoint with source_caption_id and a new style, not video_url. The docs say Sume reuses the source video of that caption job with the word timings it already has, so speech-to-text does not run again. The price does not change, because a restyle is still a render. Each accepted standalone caption job is a flat 0.20 dollars for videos up to 60 seconds.

The flow

  • Render the avatar clip with captions off so you have a clean public MP4.
  • Create a caption job with video_url and a first style. Keep the caption id from the response.
  • To try a second look, send { "source_caption_id": "...", "style": "punch" }.
  • If the transcript has a wrong word, send words with it to correct the text only.
  • Compare the results in the destination app, then keep one.

What it costs to try three looks

Three caption renders cost three flat charges. Compare that with re-rendering the avatar clip, which is billed per second, and the point is clear: caption experiments are cheap, avatar re-rolls are not.

Trying three caption styles on one 30-second plus-tier avatar clip (Sume rate card and docs, read 2026-10-07)
StepCost
Avatar clip, 30 s at plus, no product image$7.35
First caption render$0.20
Restyle 1 via source_caption_id$0.20
Restyle 2 via source_caption_id$0.20
Total$7.95

Pitfalls

Do not send video_url and source_caption_id together; the docs describe them as alternatives. If you send script_text for alignment and the clip's speech does not match, the job can fail with script_alignment_mismatch or script_alignment_failed, and the docs recommend simplifying the text or omitting it. Only one of script_text, words, cues and segments is accepted per request.

punch and tiktok-green do not support design overrides, so a restyle to those looks cannot be tuned further. Korean text on slam, punch or tiktok-green returns 400 caption_hangul_text_latin_style; this is a design guard, not a bug.

When inline captions are enough

If you need one standard look and never restyle, inline captions on the avatar request are simpler: one call, no second job. Their cost sits in the avatar estimate, they work to 60 seconds, and a failure in the caption stage is soft, leaving a clean video with captions.status=failed. Pick the standalone route when you want control, and the inline one when you want fewer steps.

Keeping the record straight

Store the first caption id next to the avatar-video id, and store each restyle's id too. When a stakeholder asks for 'the one with the yellow word', you can find it without rendering again. Use the style name in your own label, because the raw ids are opaque. Archive the rejected looks, since they cost money to make.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume