Kling 3.0 prompt guide: length, shots, dialogue, negatives
A Kling 3.0 prompt can run to 3,072 characters and include negative wording. Kling's rules for shot lists, dialogue, languages and Elements.

A Kling 3.0 prompt can be up to 3,072 characters, and Kling recommends 2,500 or fewer. It may include both positive and negative descriptions. Describe the scene, the camera and any dialogue in plain sentences; for a multi-shot clip, write each shot as “shot n, m, words;” with 1–6 shots whose seconds add up to the clip length.
The rules below come from Kling's own API reference and its VIDEO 3.0 user guide, read on 2026-09-28 and listed under Sources. The last section covers sending the same prompt to Kling 3.0 through Sume, from Video generation.
What are Kling 3.0's prompt limits?
Kling's text-to-video endpoint takes one prompt string plus settings and options; there is no separate negative-prompt field in that body (text to video).
| Rule | Kling 3.0 |
|---|---|
| Prompt length | Up to 3,072 characters; 2,500 or fewer recommended |
| Negative wording | “The prompt can include positive and negative descriptions” |
| Multi-shot | shot n, m, words; for 1–6 shots whose seconds add up to duration (3–15 s); up to 512 characters per shot |
| Elements (image to video) | Up to 3, named in the prompt as @name |
How do I write a multi-shot prompt?
Number each shot, give its length in seconds, then its prompt. For a 10-second clip, three shots of 3, 4 and 3 seconds fit. Kling's user guide describes the same choice in its app: with Multi-Shot on, the model plans shot changes itself and “will generally follow the prompts”, but adjusts when a scene suits a single shot; with Custom Multi-Shot, where you set each shot, it “will strictly follow the prompts” (user guide). Multi-shot video generation covers the settings.multi_shot switch and multi-shot on other models.
shot 1, 3, wide shot of a cyclist cresting a hill at dawn, handheld;
shot 2, 4, close-up of her hands on the handlebars, tracking;
shot 3, 3, low angle as she rides past the camera into the sun;How do I prompt dialogue, languages and accents?
Kling's user guide gives these rules for native audio (user guide):
- Pair each character with their own line; the guide's example reads “Mom (softly, in a surprised tone): Wow, I didn't expect this plot at all.”
- Dialogue can be in Chinese, English, Japanese, Korean or Spanish, mixed within one clip. Lines in any other language are translated into English.
- Tag a dialect or accent on the line, for example an Indian English accent, to have it performed.
- If an Element already has a bound voice, don't set the voice again in the prompt.
- Native audio is off by default on the API: set
settings.audiotonative.
Do negative prompts work in Kling 3.0?
Kling's API says only that the prompt “can include positive and negative descriptions” (text to video); it does not say how the model weighs them, so test “no text, no logos”-style wording on a short clip before relying on it. AI video negative prompt covers the general technique. Through Sume, POST /v1/videos has no negative-prompt field. In current code, Sume sends kling-3 a fixed negative prompt of “blur, distort, and low quality” with every request.
How do I use a Kling 3.0 prompt on Sume?
Send it as prompt with model: "kling-3" on POST /v1/videos, plus duration, aspect_ratio and generate_audio; frame_images sets a first or last frame (Video generation). In Sume's current catalog code, kling-3 takes 4–15 seconds at 16:9, 9:16 or 1:1 and no reference inputs. Sume's docs don't describe what a “shot 1, …” prompt does on kling-3, so check a short clip before building on it. Kling 3.0 API covers the rest of the request.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "kling-3",
"prompt": "Close-up of a barista pouring latte art, slow push-in, warm morning light",
"duration": 8,
"aspect_ratio": "9:16",
"generate_audio": true
}'Sources
Related posts
More in Models
- Kling 3.0 vs Kling 3.0 Omni: inputs, limits and price
Kling 3.0 and 3.0 Omni share 3–15 s clips up to 4K with native audio. Omni adds reference and edit videos; 3.0 adds Motion Control. Prices compared.
- Kling 3.0 vs Seedance 2.0: inputs, limits and price
On Sume, Kling 3.0 and Seedance 2.0 both make 4–15 s clips from text or frames. Seedance adds references and six ratios, and bills per token.
- Kling vs Wan: Kling 3.0 and Wan 3.0 by API compared
Kling 3.0 makes 4–15 s clips with no references; Wan 3.0 makes 2–30 s clips and takes image, video and audio references. Inputs and price compared.
- Kling vs Grok Imagine: Kling 3.0 and Grok Imagine Video 1.5
Kling 3.0 makes clips from text or an image, with an end frame and sound; Grok Imagine Video 1.5 animates one image, silent, for $0.0125/s.
Written by Sume