Kling O1 vs Kling 3.0 Omni: 10 to 15 seconds, audio and multi-shot
What changed from Kling Video O1 to 3.0 Omni per Kling's pages: 15 seconds, native audio, six cuts. And which of it a kling-3 request on Sume can use.

Kling VIDEO 3.0 Omni raises the clip length from 10 to 15 seconds and adds native audio and multi-shot generation, which Kling says its previous Omni model, Kling Video O1, did not have. Kling's guide describes selectable lengths of 3 to 15 seconds and up to six camera cuts in a single generation. Sume lists a kling-3 row at up to 15 seconds, but not an Omni row, so only part of this carries over.
If you built prompts for O1, the practical changes are longer scenes, spoken dialogue in the same render, and shot-by-shot direction inside one prompt. Below is what Kling's own pages say, followed by what you can actually request through Sume today.
What did Kling change from O1 to 3.0 Omni?
Kling's Omni guide compares itself to Kling Video O1 directly: the previous version was limited to 10 seconds, without native audio or multi-shot support, and 3.0 Omni adds all three. The user guide for Omni repeats the length point (up to 15 seconds, increased from 10) and lists 1080p and 720p output modes.
The guide also introduces an "AI Director" that generates up to six camera cuts in one generation, and Character Identity 3.0, which takes reference videos of 3 to 8 seconds or several images to carry a character's look, movement and voice across shots.
| Capability | Kling Video O1 | Kling VIDEO 3.0 Omni |
|---|---|---|
| Maximum length | 10 seconds | 15 seconds (selectable 3 to 15) |
| Native audio | Not available | Yes, generated in the same pass |
| Multi-shot | Not available | Up to six camera cuts per generation |
| Output modes | Not stated on the pages read | 1080p and 720p |
| Character references | Not stated on the pages read | Video of 3 to 8 seconds or several images |
What does the Omni guide say about writing shots?
Kling's user guide gives a storyboard format that states duration, framing and camera movement per shot inside the 15-second window, with references written as @ names. Its example reads like a script: "Shot 1 (2s): @Boxer A and @Boxer B face off", followed by a second shot with its own timing.
That syntax belongs to Kling's own app and its element system. The @ names point at elements you create there, so the example cannot be pasted into a Sume request and expected to bind the same elements.
What can you request through Sume?
Sume's catalog lists kling-3 with a 4 to 15 second range, 1080p accepted, and text or start and end frame inputs. The API rejects reference image and video fields on this row, and the Sume docs do not describe an Omni row or element binding. So on Sume you get the longer length and, where the row's generate_audio is true, sound, but not Kling's @ element references.
For shot-by-shot work, Sume's practical pattern is one clip per shot: generate each shot as its own 4 to 15 second job, using the previous clip's last frame as the next first frame, and join them afterwards. That is slower than six cuts in one generation, but each shot can be redone alone.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: boxing-shot-1" \
-d '{
"model": "kling-3",
"prompt": "Two boxers face off in a ring, slow push-in, crowd murmur",
"duration": 5,
"resolution": "720p",
"aspect_ratio": "16:9"
}'Should you move an O1 workflow to kling-3 on Sume?
Move it if your O1 prompts were plain text to video or image to video, because those map directly and you gain length. Stay with Kling's app, or plan a different design, if you depend on @ element references, uploaded character videos or voice binding, because those are Omni features Sume does not expose on kling-3.
Before you commit, call GET /v1/videos/models and read the live kling-3 row, then render one 5 second draft at 720p. Pricing is the provider list times 1.25 on every Sume model, and the finished job's usage.cost shows the billable amount.
Sources
Related posts
More in Comparisons
- Kling @Element character binding vs Sume: what to use instead
Kling v3 Pro on fal binds characters with elements. Sume's kling-3 takes frames only, so use a reference-capable model or one fixed first frame.
- Kling v3 Pro multi_prompt shots on fal vs Sume's one clip per shot
fal's Kling v3 Pro has a multi_prompt list for multi-shot video. Sume's kling-3 sends one prompt and frames per clip, so you join shots with Timeline 1.0.
- Kling v3 Turbo Pro at $0.14 a second vs v3 Pro and Sume's kling-3
fal lists Kling v3 Turbo Pro at $0.14 a second and v3 Pro at $0.112 audio off, $0.168 on. What that means for 5 seconds, and the Kling id Sume lists.
- Leonardo.Ai API alternative for image generation: Sume Images
Leonardo's API submits to /generations and returns a generationId. Sume POST /v1/images works from a model catalog with idempotent retries. Compared.
Written by Sume