Caption once after stitching: four 15 s clips cost $0.30, not $0.90
The caption job is $0.20 per accepted job up to 60 s. Stitch four 15 s clips first ($0.10), then caption once ($0.20): $0.30, not $0.90.

Burn captions once, on the stitched Reel. Four 15-second clips captioned one by one cost 4 x $0.20 = $0.80 and still need a $0.10 Timeline render to join them, $0.90 in all. Stitch first and caption the 60-second result and the bill is $0.10 + $0.20 = $0.30.
First Draft on an iPhone assembles your selected clips into one starting Reel (read 2026-10-09); the same order of work applies in the API: assemble, then caption the finished cut.
The two orders side by side
The caption price is flat per accepted job and is stated for videos of up to 60 seconds, so the saving only holds while the stitched Reel stays inside that figure. Timeline bills $0.10 per ceil output minute, so 60 s is one minute.
| Order | Jobs | Math | Total |
|---|---|---|---|
| Caption each clip, then Timeline | 4 captions + 1 render | 4 x $0.20 + $0.10 | $0.90 |
| Timeline, then caption once | 1 render + 1 caption | $0.10 + $0.20 | $0.30 |
| Saved | $0.60 |
Both steps in Python
The helper posts a job, polls /v1/jobs/:id/status until it is terminal and returns the result. The render result carries video_url; the caption request takes it as video_url. script_text is optional; leaving it out burns the speech-to-text words.
import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def run(path, body, key):
r = requests.post(API + path, json=body, timeout=60,
headers={**H, "Idempotency-Key": key})
r.raise_for_status()
job = r.json()["data"]["request_id"]
while True:
st = requests.get(f"{API}/v1/jobs/{job}/status", headers=H, timeout=30)
st = st.json()["data"]
if st["sume_status"] in ("completed", "failed", "canceled"):
break
time.sleep(st.get("next_poll_after_seconds") or 3)
if st["sume_status"] != "completed":
raise RuntimeError(f"{path} ended {st['sume_status']}")
res = requests.get(f"{API}/v1/jobs/{job}/result", headers=H, timeout=30)
return res.json()["data"]["result"]
cut = run("/v1/timeline-1.0/render", timeline_body, "reel-42-render")
final = run("/v1/video-captions",
{"video_url": cut["video_url"], "style": "punch"}, "reel-42-caps")Caveats
- Caption
words,cuesandsegmentstimings are capped at 60 s in the schema, matching the 60-second price note. For a longer Reel, plan one caption job per 60 s. - A stitched render with no speech in its audio spine fails with
caption_no_speech; sendcuesinstead. punchandtiktok-greenignoredesignoverrides; pickslamor a Hangul style when you need to tune colors.- Keep the same Idempotency-Key on a retry so the caption job is not billed twice.
Sources
Related posts
More in Use cases
- 45 seconds on Wan 3.0: a 30 s and a 15 s job, $11.25 at 1080p
Past the 30-second cap, a 45-second Wan 3.0 video is two jobs: 30 s and 15 s cost $11.25 at 1080p or $5.625 at 720p on Sume. Chain with the last frame.
- Character turnaround sheet on a 4:1 strip, sliced into references
One 4:1 Nano Banana 2.1 image on Sume as a four-view character sheet at $0.15 (2K), sliced into four squares and fed back as input_references. Steps and cost.
- Check a 30-second ad in 15 stills: video inspect fps 0.5 costs nothing
Review a 30-second ad as 15 stills with Sume video inspect fps 0.5. Probe and stills are free; a transcript adds $0.01 per audio minute.
- Claude Dashboards and Motion or a Sume schedule for a weekly recap
For a weekly metrics recap, Claude Dashboards plus Motion fits warehouse data and edits by hand; a Sume schedule fits footage and a $1.00 default cap per run.
Written by Sume