Mix Seedance, Kling and Omni clips in one video: shared aspect ratio

16:9 and 9:16 are the aspect ratios Seedance 2.5, Kling 3 and Gemini Omni Flash 1.1 all list. Set the Timeline output to match and plan before render.

5 min readSume
All posts

To cut Seedance, Kling and Gemini Omni Flash clips into one video on Sume, generate every shot at 16:9 (or every shot at 9:16), because that is the one pair of aspect ratios all three catalog rows list. Then set Timeline 1.0's output.width and output.height to match, and run the unbilled POST /v1/timeline-1.0/plan before you pay for the render. Do this and the fit setting has nothing to correct.

Ratios and limits are from Sume's Video generation docs and Timeline 1.0 docs, read on 2026-10-03.

Which aspect ratios do the three models share?

Seedance 2.5 lists 21:9, 16:9, 4:3, 1:1, 3:4 and 9:16, plus auto and adaptive. Kling 3 lists 16:9, 9:16 and 1:1. Gemini Omni Flash 1.1 lists 16:9 and 9:16 only. The overlap is 16:9 and 9:16, so those are the safe choice for a mixed reel. Square only works if you leave Omni out, and 21:9 only works for Seedance.

Wan 3.0 and the other rows have their own lists; read them from GET /v1/videos/models before adding them to the mix.

Aspect ratios by model for a mixed Timeline, from Sume docs, read 2026-10-03
Model idAspect ratios listedIn a 16:9 or 9:16 reel?
seedance-2.521:9, 16:9, 4:3, 1:1, 3:4, 9:16, auto, adaptiveYes
kling-316:9, 9:16, 1:1Yes
gemini-omni-flash-1.116:9, 9:16Yes

What is the Timeline default, and how do I change it?

The Timeline's default output is 1080 by 1920, a vertical frame. If your clips are 16:9 and you leave output out, they are fitted into a vertical canvas; fit is cover by default and the other values are contain, stretch and blur. Set output.width and output.height yourself, even integers between 256 and 2160, for example 1920 by 1080 for a landscape reel. Each slot needs a Sume-hosted source_url, so import clips that did not come from a Sume job first.

The same applies to frame rate. Without output.fps, the render follows the sources, and a rate that differs from a source's is met by repeating or dropping frames. Sume reports this as output_fps_resamples_sources. Different models may deliver different rates, so pick one of 24, 25, 30 or 60 on purpose.

curl -X POST https://api.sume.com/v1/timeline-1.0/plan \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "audio": {"mode": "silence", "duration_seconds": 18},
    "output": {"width": 1920, "height": 1080, "fps": 24},
    "video": [
      {"source_url": "https://media.sume.com/artifacts/artf_demo/seedance.mp4", "start": 0, "duration": 6},
      {"source_url": "https://media.sume.com/artifacts/artf_demo/kling.mp4", "start": 6, "duration": 6,
       "transition": {"type": "fade", "duration": 0.25}},
      {"source_url": "https://media.sume.com/artifacts/artf_demo/omni.mp4", "start": 12, "duration": 6,
       "transition": {"type": "fade", "duration": 0.25}}
    ]
  }'

What does the plan check, and what does it miss?

The plan runs schema checks, Sume-host URL checks and the compiler, then returns duration_seconds, segment_count, billable_minutes and estimated_cost_usd_micros. It creates no job, reserves no credits and downloads no media, and it needs no Idempotency-Key. Because it never opens the files, it cannot predict warnings about short sources that get padded or looped. Those appear on the finished render's warnings[] and are soft, not failures.

  • Generate every shot at the same ratio before assembling.
  • Set output.width, output.height and output.fps explicitly.
  • Plan first; it is free and catches schema and URL mistakes.
  • Check warnings[] on the render for padded or looped clips.
  • Remember the audio: use audio.mode: "silence" only for a silent draft, and a real spine for the final.

What does the render cost?

The render is $0.10 per output minute, rounded up, with no provider inference, so the clips' own generation costs are separate and the Timeline adds one ceil minute for a reel under 60 seconds. Confirm the live rate in GET /v1/catalog. If a single shot disappoints, the stored post re-roll one shot of a stitched AI film shows swapping one source_url and rendering again, and merge two videos of different resolutions covers the sizing side in general. To decide the models themselves, see which Sume video model for the inputs you have.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume