Chain a 30-second render, trim and captions: three jobs, three keys

Generate, trim and caption an AI clip through the Sume API as three separate jobs, each with its own Idempotency-Key, status poll and result read.

5 min readSume
All posts

To chain generate, trim and captions through the Sume API, treat each step as its own job. Submit the step, store the job id, poll GET /v1/jobs/:id/status until terminal is true, read GET /v1/jobs/:id/result, then pass the new media.sume.com URL to the next step. Send a different Idempotency-Key on each step, because each step is a different operation.

The chain matters more now that clips are longer. ByteDance describes Seedance 2.5 as generating up to 30 seconds per pass (read 2026-10-05), and the Wan 3.0 README lists native 30-second generation (read 2026-10-05). Sume's Video Router docs list seedance-2.5 at 4 to 30 seconds and wan-3.0 at 2 to 30 seconds. A 30-second render is rarely the file you publish, so the next calls are usually a cut and a caption burn.

The three steps

Each step has its own endpoint and result. There is no single call that runs all three, and there is no GET /v1/video-trim/:id. For trim, the docs say to poll the job envelope instead.

The three jobs in the chain (read 2026-10-05)
StepEndpointInputWhat you read from the result
1. GeneratePOST /v1/video-router/generatemodel, prompt, duration, resolution, aspect_ratio, mode: "async"A completed job with artifacts whose url is on media.sume.com
2. TrimPOST /v1/video-trimvideo_url, start, and one of end or durationkind: video_trim with a new video_url (a new artf_)
3. CaptionsPOST /v1/video-captionsvideo_url, optional style, script_textThe captioned video_url and artifacts

Keys: one per step, stable per retry

Use Idempotency-Key on every paid submit that your code can retry. If the same key comes back with the same operation, the retry returns the original job. If you reuse a key for a different payload, the API answers 409 idempotency_conflict. So a key such as order-8823-generate, order-8823-trim and order-8823-captions is safe: it is stable when step 2 retries, and it never collides with step 3.

If you change a step on purpose, for example a new start for the trim, that is a new operation and needs a new key. The docs say to reuse a key only for an exact retry.

Steps 2 and 3, copied from the docs

These are the documented request shapes. Replace the example video_url values with the URL from the previous step's result. Trim takes a media.sume.com artifact of your workspace; captions takes a public HTTPS video URL.

curl -X POST https://api.sume.com/v1/video-trim \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: video-trim-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
    "start": 2,
    "duration": 8
  }'

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: video-caption-script-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/example/clean.mp4",
    "style": "punch",
    "script_text": "Say hello to the Sume developer platform."
  }'

Waiting between steps

Submit with the default async mode and poll. sync waits at most 30 seconds on the HTTP request, which is a wait budget and not a job duration, so a 30-second clip will almost always return a job id with sync.timed_out true. Honor next_poll_after_seconds when it is present and use exponential backoff when it is not. Do not submit a step again only because your local worker timed out.

Read the failure from the job record, not from /result. GET /v1/jobs/:id/result answers 409 job_not_completed for jobs that are not completed, including failed ones. If the captions step fails with caption_no_speech, the clip has no audible speech: send cues instead of relying on speech-to-text.

What to store between the three steps

The chain is only as durable as the ids you keep. After each submit, write the job id, the Idempotency-Key you used and the step name to your own storage before you do anything else. If the process dies after the write, a restart reads the row and polls the existing job. If it dies before the write, a resend with the same key returns the original job rather than creating a second one.

Poll with the values the envelope gives you, not with a fixed sleep: next_poll_after_seconds says when to look again, terminal says whether the job is done, and result_ready says whether GET /v1/jobs/:id/result will answer. Reading the result on a job that is not completed returns 409 job_not_completed, so gate that call on result_ready.

Each step has its own failure surface. A render can fail for provider reasons, a trim can refuse a request with a stable code such as unsupported_media_source, and captions can fail with caption_no_speech. Handle each step's failure where it happens and do not re-run the steps before it.

  • One key per step, never shared between steps.
  • Store job id, key and step name before polling.
  • Gate the result read on result_ready.
  • Take the next step's input from the previous job's result, not from your own copy of the URL.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume