Stitching AI video clips in Sume Timeline: the rules that return 400

Timeline 1.0 joins model clips into one MP4: first slot at 0, import first, fades up to 1 s, at most 8 chained fades. Error codes, cost, and a curl request.

5 min readSume
All posts

To join AI video clips on Sume, import each clip to your media store, then send Timeline 1.0 an audio spine plus ordered video[] slots. The first slot must start at 0, transitions go only on later slots, a fade is at most 1 s, and more than 8 chained fades is refused. Rendering costs $0.10 per output minute, rounded up.

Get clips onto the media host

Every source_url must already be an artifact or asset of your workspace on media.sume.com. Off-host links fail with unsupported_media_source. Import files with POST /v1/media-imports first, and keep the returned URL.

The slot rules

The Timeline docs list the stable error codes. These are the ones that trip people who stitch Wan, Seedance and Omni clips of different lengths.

From apps/docs/content/docs/models/timeline.md, read 2026-10-05
CodeCauseFix
timeline_must_start_at_zerovideo[0].start is not 0Start the first slot at 0
transition_on_first_segmentTransition on video[0]Move it to the second slot
segment_overlap / invalid_segment_timingStarts do not increase, or slots overlap past the crossfadeMake each start the previous start plus its duration
transition_too_longOver 1 s or over 50% of the shorter neighborShorten the fade
too_many_chained_transitionsMore than 8 adjacent fadesInsert a hard cut
render_strategy_unsafestrategy single above 12 slotsUse auto or chunked

Cost of a long join

The reserve is ceil(audio.duration_seconds / 60) minutes at $0.10 each. Three 30 s clips make 90 s, so the render is 2 minutes and $0.20. The clips themselves cost far more: three Wan 3.0 clips at 720p are 90 s x $0.125 = $11.25. The Timeline job uses worker ffmpeg only, with no provider inference. Use the unbilled plan call to check the timing first, as described in our plan note.

A request

The body below follows the example in the docs: one voiceover spine, two clips, a 0.25 s fade on the second. Idempotency-Key is required. Without mode: "sync" you get 202 and poll /v1/jobs/:id/status.

curl -X POST https://api.sume.com/v1/timeline-1.0/render \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: stitch-001" \
  -d '{
    "audio": { "url": "https://media.sume.com/artifacts/artf_demo/voice.wav", "duration_seconds": 24 },
    "video": [
      { "source_url": "https://media.sume.com/artifacts/artf_demo/intro.mp4", "start": 0, "duration": 8 },
      { "source_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4", "start": 8, "duration": 16,
        "transition": { "type": "fade", "duration": 0.25 } }
    ]
  }'

Frame rate

If you omit output.fps, the render follows the sources. Mixed rates (an Omni clip at 24 fps beside a 30 fps clip) can resample and judder, and the job reports output_fps_resamples_sources. Set one rate from 24, 25, 30 or 60 when you mix models. The default output is 1080x1920, so set output.width and output.height for landscape.

Preflight with plan

POST /v1/timeline-1.0/plan runs the schema, the host checks and the compiler, and returns duration_seconds, segment_count, billable_minutes and estimated_cost_usd_micros without creating a job. It cannot predict warnings about short sources that get padded or looped. For a dozen clips, run plan first, fix the codes in the table, then render with an Idempotency-Key.

Past 12 segments the default strategy chunks the render. That is automatic and needs no change on your side, but single is refused above 12 slots.

Related posts

More in Media tools

All Media tools posts

Written by Sume