Restyle one captioned clip into three looks with source_caption_id

Pass source_caption_id to video-captions to re-burn the same clip in a new style without a second speech-to-text run. A restyle is still a billed render.

4 min readSume
All posts

Send source_caption_id instead of video_url to POST /v1/video-captions. Sume reuses the first job's source video and its word timings, so no second speech-to-text runs, and you pick a different style or design. The docs are explicit that billing is unchanged: a restyle is still a render, so three looks is three renders, not one.

What changes and what does not

From the video captions docs: pass words alongside the id only to correct the wording. style accepts slam, punch, tiktok-green and others; design overrides colours, typography, placement, phrasing and motion, but is not supported on punch or tiktok-green.

Restyle behaviour (Sume docs, read 2026-10-04)
ItemRestyle with source_caption_id
Source videoReused
Word timingsReused, no second STT
BillingUnchanged, still a render
design on punch or tiktok-greenNot supported

Three looks

Keep the first job id, then submit a new job per look with a distinct Idempotency-Key. The docs list a fixed $0.20 per accepted standalone caption job for videos up to 60 seconds, so three looks of a short clip come to about $0.60 (confirm in GET /v1/catalog). Poll each job as described in Jobs and results.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: restyle-slam-001" \
  -d '{
    "source_caption_id": "vc_123",
    "style": "slam",
    "design": { "placement": { "anchor_ratio": 0.55 } }
  }'

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume