Restyle captions three times with source_caption_id: $0.60 total
To test three caption looks on one video, send source_caption_id: no second transcription, but each restyle is a render at the same job price.

To try a second caption style on a video that Sume has already captioned, send source_caption_id instead of video_url. Sume reuses that caption's source video and the word timings it already holds, so speech-to-text does not run again. The price does not change, because a restyle is still a render. Three looks on one clip is three jobs at $0.20 each, $0.60 in total, under the docs' current estimate for clips up to 60 seconds.
The arithmetic
The caption docs state each accepted standalone job reserves and captures $0.20, and that a restyle is priced like any other job. The live price is in GET /v1/catalog.
| Job | Input | Price |
|---|---|---|
| 1. First caption | video_url, style black-outline | $0.20 |
| 2. Restyle | source_caption_id, style weight-shift | $0.20 |
| 3. Restyle | source_caption_id, style clip-wipe | $0.20 |
| Total | 3 jobs | 3 x $0.20 = $0.60 |
What restyling saves
The saving is not money; it is consistency and time. The transcript and timings are fixed after the first job, so the three renders differ only in look. If a word is wrong, send words along with source_caption_id to correct the text.
The same applies to a typo you fix once: correct it in the restyle request, then keep the corrected job as your base for later restyles.
Request
Replace the id with the one from your first caption job.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: restyle-clip-wipe-001" \
-d '{
"source_caption_id": "vc_123",
"style": "clip-wipe"
}'Picking between Hangul looks
For Korean speech the docs describe clip-wipe as the clearest style at small phone sizes, and black-outline as the safe default. highlight moves an accent block behind the spoken word and pill-karaoke changes color inside a dark pill. Compare them on one real clip, because the way a style reads depends on the footage behind it.
When not to restyle
Restyle works from a finished caption job. If the transcript itself is wrong, change the words rather than the style. If the clip changed, upload the new video as a new job, because source_caption_id reuses the original source video and its timings.
Keep each caption job id next to the video it came from. Then a later brand refresh costs $0.20 per clip for a render, with no re-transcription and no new upload. Forty clips refreshed would be 40 x $0.20 = $8.00 under the same estimate.
- Standalone caption job: $0.20 under the current estimate for clips up to 60 seconds.
- Inline captions on an Avatar Video are a separate add-on and create no caption resource, so they cannot be restyled this way.
- Confirm the live rate in
GET /v1/catalogbefore a large batch.
Sources
Related posts
More in Media tools
- script_alignment_mismatch: fixing a caption job with script_text
When script_text does not match the speech, a Sume caption job fails with script_alignment_mismatch or script_alignment_failed. Simplify the script or omit it.
- Six 5-second reference clips: $2.28 on H3 to $17.34 on Seedance 2.5
Cost of six 5-second 720p reference-to-video clips on Sume by model: MiniMax H3 $2.28 at 768p, Wan 3.0 $3.78, Seedance 2.0 Fast $9.12, Seedance 2.5 $17.34.
- Timeline soundtrack duck_db: 0 to 20 dB, needs a real audio spine
Sume Timeline soundtrack.duck_db accepts 0 to 20 and returns duck_requires_audio_spine on silence. Use it with a voiceover spine; gain_db covers a static level.
- STT-ready audio: detach as 16000 Hz mono wav in one call
Sume audio detach can output the STT shape directly: sample_rate 16000 with channels mono, in sample-exact wav, for $0.01 per job.
Written by Sume