Kling 3.0 starts at 3 seconds; Sume's kling-3 starts at 4
Kling's own page lists 3-15 seconds for Kling 3.0. The kling-3 id on Sume accepts 4-15, so a 3 second request returns unsupported_capability.

A 3 second clip works on Kling's own product but not on the kling-3 id on Sume: Kling's comparison page lists Kling 3.0 at 3-15 seconds, while the Sume catalog gives kling-3 ('Kling Video v3 Pro') a 4-15 second range. Send duration: 3 and Sume answers HTTP 400 with unsupported_capability and a message of the form 'kling-3 does not support duration 3', plus the list of accepted values. Send 4 and it runs.
Kling's figures come from Kling 4.0 vs 3.0 on kling.ai, read on 2026-10-03. The Sume side comes from the catalog in the Video Router docs and Video generation docs. Sume does not clamp or round a duration for you on a pinned id; it rejects it, so nothing is billed for a request that cannot run.
Where the two ranges differ
The mismatch is small but it bites the shortest hook clips, which are the ones people most often generate at 3 seconds. The table shows the same field on each side.
Kling 4.0 is a separate case: the same Kling page lists 3-30 seconds for it, but Kling 4.0 is not an id in the Sume catalog today, so there is nothing to call. A pinned request for it fails on the model id before duration is even checked.
| Model | Kling page range | Sume id | Sume range |
|---|---|---|---|
| Kling 3.0 | 3-15 s | kling-3 | 4-15 s |
| Kling 4.0 | 3-30 s | not in the catalog | not callable |
| Seedance 2.5 | up to 30 s per generation | seedance-2.5 | 4-30 s |
| Wan 3.0 | not read | wan-3.0 | 2-30 s |
| Gemini Omni Flash 1.1 | not read | gemini-omni-flash-1.1 | 3-10 s |
What to do for a 3 second slot
You have three honest options. Round up to 4 seconds on kling-3 and trim the extra second afterwards with the video trim tool, which cuts an existing clip to an exact length. Or pin a model whose floor is lower: wan-3.0 starts at 2 seconds and gemini-omni-flash-1.1 at 3. Or keep Kling and plan the edit around a 4 second beat.
Trimming costs nothing in generation terms, but you pay for the 4 seconds you generated, at list price times 1.25 like every Video Router model. If the slot is a fixed 3 seconds in a template, the cheaper path is usually a model that accepts 3 natively, because the extra second is wasted spend multiplied by every variant you run.
Read the range instead of hard-coding it
Do not copy the numbers from this page into code. GET /v1/videos/models returns supported_durations for every id, and the Video Router catalog returns the same envelope under capabilities.duration_seconds. Build the request from that list, and a model whose range changes later will not break your job.
The same habit protects you from vendor marketing pages, which describe the vendor's product rather than the id behind a third-party router. When a vendor page and the Sume catalog disagree, the catalog is what your request is checked against.
Sources
Related posts
More in Models
- Kyutai Pocket TTS languages: what it speaks vs hosted Sume TTS
Pocket TTS is a 100M-parameter open model you run yourself. Its README lists seven languages; Sume's TTS is hosted, per character, with a voice id.
- LLM knowledge cutoffs: Claude Jun 2026, GPT-6.1 Sol Apr 2026
Claude 5.5 models list a June 2026 cutoff and GPT-6.1 Sol April 30, 2026. A video agent on trends needs a data tool; Sume ships trending-videos search.
- Longest AI video clip in one request: 30, 15 or 10 seconds by model
Seedance 2.5 and Wan 3.0 reach 30 seconds on Sume; most other rows stop at 15 and Gemini Omni Flash at 10. Ceilings per model, and when to stitch instead.
- LTX-2 diffusion decoder or convolutional decoder: which to use
LTX-2 ships a diffusion decoder (better quality, more VRAM) and a lighter convolutional one. What the README says, a draft-then-final habit, and hosted jobs.
Written by Sume