AI video prompt JSON: what the keys actually do
A JSON prompt for AI video is still text: the model reads keys as words. Length, size and audio are request fields, not keys inside the prompt.

A JSON prompt for AI video is still a text prompt: the video model receives your keys and values as one string of words, not as settings. Clip length, resolution, aspect ratio, and audio are set by separate request fields, and a JSON prompt that lists five scenes is still one generation with one clip length.
This post uses Sume's Video generation API as the worked example, plus Agent Completions and Structured output for JSON shot plans, all read on 2026-09-28. "Current code" marks behavior read from Sume's API code rather than the docs.
Does the API read the keys in a JSON prompt?
No. In the API contract, prompt is a string: a text description of the video to generate. Braces, quotes, and key names are characters inside that string. For a pinned catalog model, Sume passes the prompt on to that model; in current code the kling-3 request, for example, sends your prompt text as its prompt, while clip length and audio come from the duration and generate_audio fields.
The same code also sends kling-3 a fixed negative prompt, "blur, distort, and low quality", so a negative key you write inside the prompt is just more prompt text.
curl -X POST "https://api.sume.com/v1/videos" \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "kling-3",
"prompt": "{\"scene\": \"barista pours latte art\", \"camera\": \"slow push-in\", \"lighting\": \"warm morning window light\"}",
"duration": 10,
"resolution": "720p",
"aspect_ratio": "9:16",
"generate_audio": false
}'Where do the usual JSON prompt keys go instead?
Keys that describe the picture can stay in the prompt text. Keys that set the output belong in request fields; How to write an AI video prompt has notes on each field.
| Key in a JSON prompt | Where it goes on Sume |
|---|---|
| scene, subject, camera, lighting, motion | prompt text |
| duration, length | duration (seconds) |
| resolution, quality | resolution, such as 720p |
| aspect ratio, orientation | aspect_ratio, such as 9:16 |
| audio on or off | generate_audio |
| start or end image | frame_images: a first_frame, optionally with a last_frame (current code refuses a last_frame alone) |
| style reference | input_references |
Do JSON prompts give better AI videos?
Sume's docs make no claim either way, and neither does this post. What the docs do ask for is detail: specific, descriptive prompts with motion, camera angles, lighting, and scene composition. Labeled keys are one way to make sure each of those is written down. For prompt wording itself, see How to write a prompt for an AI video generator.
How do I turn a JSON shot plan into several clips?
One POST /v1/videos request is one job with one duration, so a multi-scene JSON prompt is still one generation, capped at that model's clip length. For a real shot list, get the plan as structured JSON first, then send each shot as its own request; Multi-shot video generation covers joining the clips.
Sume's Agent Completions take an output_schema that binds the run's output to your own JSON Schema. generation_spend_cap_usd is required and has no default:
- The create call answers
202with a receipt; poll itsstatus_urluntilnext_actionstops beingpoll_status, then readoutput.shots. - Over the API, a run whose output can't satisfy your schema is
failed, with the reason inoutput_error, so check that field before readingoutput. - Send each shot as its own
/v1/videosrequest with its owndurationandaspect_ratio. Output schema templates has more schemas to copy.
curl -sS -X POST "https://api.sume.com/v1/agent/completions" \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"instruction": "Plan a 3-shot, 15-second coffee ad. One video prompt per shot. Do not generate media.",
"output_schema": {
"name": "shot_plan",
"schema": {
"type": "object",
"properties": { "shots": { "type": "array", "items": { "type": "string" } } },
"required": ["shots"],
"additionalProperties": false
}
},
"generation_spend_cap_usd": 1
}'Sources
Related posts
More in Models
- Kling 3.0 API: text or image to video, limits and price
Kling 3.0 has an API: Kling's own, and multi-model APIs such as Sume's, where it is kling-3: 4–15 second clips from text or frames, audio optional.
- Kling 3.0 prompt guide: length, shots, dialogue, negatives
A Kling 3.0 prompt can run to 3,072 characters and include negative wording. Kling's rules for shot lists, dialogue, languages and Elements.
- Kling 3.0 vs Kling 3.0 Omni: inputs, limits and price
Kling 3.0 and 3.0 Omni share 3–15 s clips up to 4K with native audio. Omni adds reference and edit videos; 3.0 adds Motion Control. Prices compared.
- Kling 3.0 vs Seedance 2.0: inputs, limits and price
On Sume, Kling 3.0 and Seedance 2.0 both make 4–15 s clips from text or frames. Seedance adds references and six ratios, and bills per token.
Written by Sume