Kling 3.0 in the Sume Videos panel: 5-15 s choices vs API 4-15 s
The panel offers Kling 3.0 at 5, 6, 8, 10 or 15 seconds and no end frame. The API accepts any whole second from 4 to 15 and a last frame.

Which durations does the Sume panel offer for Kling 3.0?
Five: 5, 6, 8, 10 and 15 seconds. That is the list in the panel's control definition for Kling 3.0, alongside aspect ratios of 16:9, 9:16 and 1:1 and an audio toggle.
The API row is wider. The kling-3 catalog entry accepts whole seconds from 4 to 15, so 4, 7, 9, 11, 12, 13 and 14 are valid API requests that the panel does not present.
| Control | Videos panel | API row kling-3 |
|---|---|---|
| Durations | 5, 6, 8, 10, 15 s | Whole seconds, 4 to 15 |
| Aspect ratios | 16:9, 9:16, 1:1 | 16:9, 9:16, 1:1 |
| End frame | Not exposed for create | First and last frame accepted |
| Resolutions | 720p, 1080p, plus a 4K label | 720p, 1080p |
| Audio | Include-audio toggle | generate_audio boolean |
Why is there no end frame in the panel?
The panel's Kling 3.0 definition sets end-frame support to false, with a code comment noting that the product path does not expose an end or tail frame for create. The API does: the catalog lists end_frame: true and supported_frame_images as first and last frame, with the rule that a last frame must come with a first frame.
So a looping clip or a defined landing pose is an API job today. The first and last frame guide shows the request shape.
What about the 4K option?
The panel shows a 4K label for Kling 3.0, but the API row stops at 1080p, and a 4K request against kling-3 is rejected as not advertised. Rather than repeat the detail here, read the 4K option explainer, which covers what the panel submits.
How do I pick the length for a 7 second idea?
Either round up to 8 in the panel and trim, or call the API with duration: 7. Trimming a longer clip is cheap to reason about because the reserve on submit scales with seconds: the workspace balance is held at the provider list rate times 1.25, and the finished job reports the Sume billable amount in usage.cost.
If the exact second matters for lip-sync or a beat in the music, request it directly through the API instead of trimming.
curl -X POST "https://api.sume.com/v1/videos" \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "kling-3",
"prompt": "A paper boat drifts across a rainy street, slow tracking shot",
"duration": 7,
"resolution": "720p",
"aspect_ratio": "1:1"
}'How do I keep the two surfaces from drifting apart?
Treat GET /v1/videos/models as the source of truth for limits and the panel as a convenience view of it. When a number in the panel and a number in the catalog disagree, the catalog is what the API enforces. Pin kling-3 explicitly in code, because Auto does not disclose which family served a clip, and read Kling 3 duration and aspect limits for the full envelope.
Sources
Related posts
More in Models
- Kyutai Pocket TTS languages: what it speaks vs hosted Sume TTS
Pocket TTS is a 100M-parameter open model you run yourself. Its README lists seven languages; Sume's TTS is hosted, per character, with a voice id.
- Longest AI video clip in one request: 30, 15 or 10 seconds by model
Seedance 2.5 and Wan 3.0 reach 30 seconds on Sume; most other rows stop at 15 and Gemini Omni Flash at 10. Ceilings per model, and when to stitch instead.
- LTX-2 diffusion decoder or convolutional decoder: which to use
LTX-2 ships a diffusion decoder (better quality, more VRAM) and a lighter convolutional one. What the README says, a draft-then-final habit, and hosted jobs.
- LTX-2 Dub-It: lip-matched rephrasing vs Sume's dubbing steps
LTX-2's Dub-It pipeline rephrases speech while matching speaker and lips. What the README lists, what it omits, and what Sume's dubbing steps do and do not do.
Written by Sume