AI video prompt JSON: what the keys actually do

A JSON prompt for AI video is still text: the model reads keys as words. Length, size and audio are request fields, not keys inside the prompt.

5 min readSume
All posts

A JSON prompt for AI video is still a text prompt: the video model receives your keys and values as one string of words, not as settings. Clip length, resolution, aspect ratio, and audio are set by separate request fields, and a JSON prompt that lists five scenes is still one generation with one clip length.

This post uses Sume's Video generation API as the worked example, plus Agent Completions and Structured output for JSON shot plans, all read on 2026-09-28. "Current code" marks behavior read from Sume's API code rather than the docs.

Does the API read the keys in a JSON prompt?

No. In the API contract, prompt is a string: a text description of the video to generate. Braces, quotes, and key names are characters inside that string. For a pinned catalog model, Sume passes the prompt on to that model; in current code the kling-3 request, for example, sends your prompt text as its prompt, while clip length and audio come from the duration and generate_audio fields.

The same code also sends kling-3 a fixed negative prompt, "blur, distort, and low quality", so a negative key you write inside the prompt is just more prompt text.

curl -X POST "https://api.sume.com/v1/videos" \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "kling-3",
    "prompt": "{\"scene\": \"barista pours latte art\", \"camera\": \"slow push-in\", \"lighting\": \"warm morning window light\"}",
    "duration": 10,
    "resolution": "720p",
    "aspect_ratio": "9:16",
    "generate_audio": false
  }'

Where do the usual JSON prompt keys go instead?

Keys that describe the picture can stay in the prompt text. Keys that set the output belong in request fields; How to write an AI video prompt has notes on each field.

From Video generation, read 2026-09-28.
Key in a JSON promptWhere it goes on Sume
scene, subject, camera, lighting, motionprompt text
duration, lengthduration (seconds)
resolution, qualityresolution, such as 720p
aspect ratio, orientationaspect_ratio, such as 9:16
audio on or offgenerate_audio
start or end imageframe_images: a first_frame, optionally with a last_frame (current code refuses a last_frame alone)
style referenceinput_references

Do JSON prompts give better AI videos?

Sume's docs make no claim either way, and neither does this post. What the docs do ask for is detail: specific, descriptive prompts with motion, camera angles, lighting, and scene composition. Labeled keys are one way to make sure each of those is written down. For prompt wording itself, see How to write a prompt for an AI video generator.

How do I turn a JSON shot plan into several clips?

One POST /v1/videos request is one job with one duration, so a multi-scene JSON prompt is still one generation, capped at that model's clip length. For a real shot list, get the plan as structured JSON first, then send each shot as its own request; Multi-shot video generation covers joining the clips.

Sume's Agent Completions take an output_schema that binds the run's output to your own JSON Schema. generation_spend_cap_usd is required and has no default:

  • The create call answers 202 with a receipt; poll its status_url until next_action stops being poll_status, then read output.shots.
  • Over the API, a run whose output can't satisfy your schema is failed, with the reason in output_error, so check that field before reading output.
  • Send each shot as its own /v1/videos request with its own duration and aspect_ratio. Output schema templates has more schemas to copy.
curl -sS -X POST "https://api.sume.com/v1/agent/completions" \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "instruction": "Plan a 3-shot, 15-second coffee ad. One video prompt per shot. Do not generate media.",
    "output_schema": {
      "name": "shot_plan",
      "schema": {
        "type": "object",
        "properties": { "shots": { "type": "array", "items": { "type": "string" } } },
        "required": ["shots"],
        "additionalProperties": false
      }
    },
    "generation_spend_cap_usd": 1
  }'

Sources

Related posts

More in Models

All Models posts

Written by Sume