Which URL goes into video-captions after a trim: the new artf_
After video-trim, caption the new video_url from the trim result, not the source render. It is a new artifact on media.sume.com and a public HTTPS URL.

Pass the trim job's own video_url to POST /v1/video-captions, not the URL of the original render. The trim result returns a new artifact (artf_) and the docs say it is never the source. If you caption the source by mistake, the burned captions will cover the whole render and your cut will not be in the output.
Reading the right field
When the trim job's status is result_ready, GET /v1/jobs/:id/result returns kind: video_trim with video_url, duration_seconds, actual_start_seconds, precision, audio, output and optional warnings[]. Take video_url from there. There is no GET /v1/video-trim/:id, so the job result is the only read path.
| Step | Field you read | Field you set in the next call |
|---|---|---|
| Trim result | video_url (new artf_) | video_url on POST /v1/video-captions |
| Trim result | actual_start_seconds | Only for your own timing, if you used keyframe |
| Caption result | Captioned video_url and artifacts | What you publish |
| Caption job | GET /v1/video-captions/:id resource | Source for source_caption_id restyles |
The caption request
The caption input must be a public HTTPS URL that Sume can fetch. The API rejects localhost, private-network, non-HTTPS, signed or private URLs and provider task URLs. A media.sume.com artifact URL meets that. style, font, language, script_text, words, cues and segments are optional, and only one of script_text, words, cues and segments can be sent.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: video-caption-script-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/example/clean.mp4",
"style": "punch",
"script_text": "Say hello to the Sume developer platform."
}'Two failure modes after a trim
A trimmed range can drop the speech. If the cut has no audible speech, the caption job fails with caption_no_speech and next_action: use_overlay_captions; send cues with text, start and end in seconds measured from the trimmed clip's start. If you send script_text and the speech-to-text timing cannot be aligned to it, the job fails with script_alignment_mismatch or script_alignment_failed; simplify the script or omit it.
To try a different look later, send source_caption_id in place of video_url. Sume reuses the earlier word timings, so speech-to-text does not run again, and the price is still $0.20 because a restyle is a render.
A guard against the wrong URL
Make the chain code pass the URL by reference, not by retyping it. Keep a field such as current_video_url on your chain row. After each step completes, overwrite it with that step's output URL. The captions step reads current_video_url only. That way a code change cannot send the original render into a step that expects the trimmed cut.
Log the URL at each hop with its step name. If a viewer reports captions over the full clip, the log shows at once which URL the caption job received. The check is free and settles the question without a new run.
Keep in mind that a trim leaves its source untouched. The original render is still there, so you can trim it again with other times if needed.
- One field for the current URL.
- Overwrite it after each step.
- Log the URL with the step name.
What the final file looks like
The caption job's result carries the captioned video and its artifacts, and that is the URL to publish. Keep the trimmed clip as well: it is the clean cut, and a restyle with source_caption_id starts from the caption resource, not from re-running speech-to-text.
If the caption step fails, nothing is lost upstream. The trim result is still stored, so you can fix the cause, such as sending cues for a silent cut, and submit the caption step again with a new key and the corrected body.
Sources
Related posts
More in Developers
- Which video model takes 30 seconds? Filter the Sume catalog in Python
Ask GET /v1/videos/models which models accept duration 30 at 1080p instead of hard-coding ids. A short Python filter, and why Omni and Kling 3 drop out.
- Which video models accept 30 seconds on Sume: read the catalog first
Seedance 2.5, Wan 3.0, h3-max-recast and Genjutsu list 30 s on the Sume Video Router; MiniMax H3 stops at 15 s. Read the catalog before you set duration.
- Why a Kling motion control job shows type avatar_image_to_video
Sume stores Kling motion control as type avatar_image_to_video with model kling/3.0/motion-control. The model id, not the type, tells it apart from Fabric.
- Why Sume MCP results show [redacted]: the redaction field list
Sume MCP omits api_key fields and masks secrets and signed URLs as [redacted] in tool results. That is intended, not a failed call.
Written by Sume