Kling 3.0 multi-shot: Sume's kling-3 has no field, so cut on Timeline

Kling's guide describes automatic and custom multi-shot for Kling 3.0. Sume's kling-3 request has no multi-shot field; join separate clips with Timeline.

4 min readSume
All posts

Kling's VIDEO 3.0 guide lists multi-shot as a feature, with automatic multi-shot planning or a custom multi-shot configuration (read 2026-10-04). The page does not state a maximum shot count. If you call Kling 3.0 through Sume you should know what survives the trip.

Sume's kling-3 is Kling Video v3 Pro behind the Video Router. Its documented input shapes are text-to-video and start or end frames, with audio on or off. The request body in the video generation docs has prompt, duration, resolution, aspect_ratio, frame_images, input_references and generate_audio; none of them is a multi-shot configuration, and I found no multi-shot field for kling-3 in the API schema. Describing cuts in the prompt is allowed as text, but Sume does not expose Kling's structured multi-shot controls.

When you need hard cuts

Make each shot its own clip and assemble them. Timeline 1.0 takes ordered video[] slots plus an audio spine and returns one MP4. Each slot after the first can carry a transition (fade, wipeleft, wiperight, slideup, slidedown, dissolve), and the default output is 1080 by 1920.

Two practical limits help. A slot runs at least 0.2 seconds, and a single render holds 1 to 200 slots. The public rate is $0.10 per output minute, so a 30-second reel of three Kling shots adds $0.10 for the join, on top of the three clips.

Trade-offs

Check the live catalog before you plan: GET /v1/videos/models returns the fields each model supports, and anything absent there is not available through Sume today.

  • Separate clips cost more than one multi-shot generation only if the vendor bills multi-shot as one clip; Sume bills each kling-3 request by seconds, audio on or off.
  • Characters can drift between separate clips. Use the same first-frame image for each shot, since kling-3 takes start and end frames.
  • Audio is generated per clip. If you want one voice track, set generate_audio to false and use the Timeline audio spine instead.

Sources

Related posts

More in Models

All Models posts

Written by Sume