Mix Seedance, Kling and Omni clips in one video: shared aspect ratio
16:9 and 9:16 are the aspect ratios Seedance 2.5, Kling 3 and Gemini Omni Flash 1.1 all list. Set the Timeline output to match and plan before render.

To cut Seedance, Kling and Gemini Omni Flash clips into one video on Sume, generate every shot at 16:9 (or every shot at 9:16), because that is the one pair of aspect ratios all three catalog rows list. Then set Timeline 1.0's output.width and output.height to match, and run the unbilled POST /v1/timeline-1.0/plan before you pay for the render. Do this and the fit setting has nothing to correct.
Ratios and limits are from Sume's Video generation docs and Timeline 1.0 docs, read on 2026-10-03.
Which aspect ratios do the three models share?
Seedance 2.5 lists 21:9, 16:9, 4:3, 1:1, 3:4 and 9:16, plus auto and adaptive. Kling 3 lists 16:9, 9:16 and 1:1. Gemini Omni Flash 1.1 lists 16:9 and 9:16 only. The overlap is 16:9 and 9:16, so those are the safe choice for a mixed reel. Square only works if you leave Omni out, and 21:9 only works for Seedance.
Wan 3.0 and the other rows have their own lists; read them from GET /v1/videos/models before adding them to the mix.
| Model id | Aspect ratios listed | In a 16:9 or 9:16 reel? |
|---|---|---|
seedance-2.5 | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, auto, adaptive | Yes |
kling-3 | 16:9, 9:16, 1:1 | Yes |
gemini-omni-flash-1.1 | 16:9, 9:16 | Yes |
What is the Timeline default, and how do I change it?
The Timeline's default output is 1080 by 1920, a vertical frame. If your clips are 16:9 and you leave output out, they are fitted into a vertical canvas; fit is cover by default and the other values are contain, stretch and blur. Set output.width and output.height yourself, even integers between 256 and 2160, for example 1920 by 1080 for a landscape reel. Each slot needs a Sume-hosted source_url, so import clips that did not come from a Sume job first.
The same applies to frame rate. Without output.fps, the render follows the sources, and a rate that differs from a source's is met by repeating or dropping frames. Sume reports this as output_fps_resamples_sources. Different models may deliver different rates, so pick one of 24, 25, 30 or 60 on purpose.
curl -X POST https://api.sume.com/v1/timeline-1.0/plan \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"audio": {"mode": "silence", "duration_seconds": 18},
"output": {"width": 1920, "height": 1080, "fps": 24},
"video": [
{"source_url": "https://media.sume.com/artifacts/artf_demo/seedance.mp4", "start": 0, "duration": 6},
{"source_url": "https://media.sume.com/artifacts/artf_demo/kling.mp4", "start": 6, "duration": 6,
"transition": {"type": "fade", "duration": 0.25}},
{"source_url": "https://media.sume.com/artifacts/artf_demo/omni.mp4", "start": 12, "duration": 6,
"transition": {"type": "fade", "duration": 0.25}}
]
}'What does the plan check, and what does it miss?
The plan runs schema checks, Sume-host URL checks and the compiler, then returns duration_seconds, segment_count, billable_minutes and estimated_cost_usd_micros. It creates no job, reserves no credits and downloads no media, and it needs no Idempotency-Key. Because it never opens the files, it cannot predict warnings about short sources that get padded or looped. Those appear on the finished render's warnings[] and are soft, not failures.
- Generate every shot at the same ratio before assembling.
- Set
output.width,output.heightandoutput.fpsexplicitly. - Plan first; it is free and catches schema and URL mistakes.
- Check
warnings[]on the render for padded or looped clips. - Remember the audio: use
audio.mode: "silence"only for a silent draft, and a real spine for the final.
What does the render cost?
The render is $0.10 per output minute, rounded up, with no provider inference, so the clips' own generation costs are separate and the Timeline adds one ceil minute for a reel under 60 seconds. Confirm the live rate in GET /v1/catalog. If a single shot disappoints, the stored post re-roll one shot of a stitched AI film shows swapping one source_url and rendering again, and merge two videos of different resolutions covers the sizing side in general. To decide the models themselves, see which Sume video model for the inputs you have.
Sources
Related posts
More in Developers
- Perplexity Decisions API as a publish gate for Sume output
Check a finished Sume Format image with Perplexity's Decisions API before it ships: base64 data URL, one yes/no question, a threshold, and a human-review lane.
- Pocket TTS API: run it yourself or call a hosted TTS API
Kyutai's Pocket TTS installs with pip and serves from localhost. If you want a hosted API with job URLs instead, here is the Sume request and what changes.
- Prefect 3 task retries for a Sume job: same Idempotency-Key
Retry a Sume image job in Prefect 3 without paying twice: a tested flow with retry_condition_fn, delay list and an order-derived idempotency key.
- Pydantic model to Sume output_schema: extra forbid, no defaults
Turn a Pydantic v2 model into a valid Sume Format output_schema: extra=forbid, nullable instead of defaults, and the SumeMediaFile reference. Tested.
Written by Sume