'kling-3 supports text-to-video or start/end frames only': the fix
kling-3 on Sume has no reference lists. The 400 is fixed by sending text only, or a first frame and optional end frame, or by switching to a reference model.

The Video Router message "model kling-3 supports text-to-video or start/end frames only (no reference_*_urls)." means the request contained reference_image_urls, reference_video_urls or reference_audio_urls. For kling-3, send a prompt alone, or a prompt with image_url and optionally end_image_url. For reference-guided shots, use a model with reference support.
What kling-3 does take
In the catalog kling-3 (Kling Video v3 Pro) has text-to-video, image-to-video and an end frame, no reference images, videos or audios, resolutions of 720p and 1080p at the catalog level, and 4 to 15 seconds. Its catalog entry also carries a separate audio-on list price. One caveat: the Video Router validation lists the models that may use 1080p, and kling-3 is not among them, so check the live model endpoint for what the route accepts before you pin 1080p.
| You want | Use | Reference limits |
|---|---|---|
| Reference images | seedance-2.5, wan-3.0, minimax-h3 | Up to 9 (Seedance, MiniMax) or 10 (Wan) |
| Reference videos | seedance-2.5, wan-3.0, minimax-h3 | Up to 3 (Seedance, MiniMax) or 5 (Wan) |
| Reference audio | seedance-2.5, wan-3.0, minimax-h3 | Up to 3 or 5, with an image or video |
| Motion from a clip, same framing | higgsfield-genjutsu | One video plus 1 to 8 images |
| Only a first and last frame | kling-3 | Frames, no references |
Convert references into frames
If you can pick one image as the opening, a reference image becomes a first frame, and the request is valid for Kling. The converter does exactly that, refusing a request that has anything but images, since videos and audio have no frame equivalent. It prints the body and sends nothing.
def kling_ready(body: dict) -> dict:
out = dict(body)
if out.get("reference_video_urls") or out.get("reference_audio_urls"):
raise ValueError("kling-3 has no video or audio references; change model")
refs = out.pop("reference_image_urls", [])
if refs:
out.setdefault("image_url", refs[0])
return out
print(kling_ready({
"model": "kling-3",
"prompt": "The courier smiles and waves",
"reference_image_urls": ["https://example.com/courier.png"],
}))
When the swap is wrong
A reference image and a first frame are not the same instruction. A reference tells the model what a subject looks like; a first frame is the picture the clip starts from. If you only need the look of a character across many shots, a reference model is the right fit, and forcing it into a first frame will fix the error while changing the shot. The related post on kling-3 and motion_video_url covers the other route that takes a clip.
Before you migrate a whole workflow, run one reference-style shot on both models at the same duration and resolution. Compare how well the subject holds, and compare the per-second price in each catalog entry. The cheapest model that keeps the character recognisable is the right one; a model that is cheaper per second but needs three retries is not.
Idempotency keys also deserve a thought. A corrected request is a new body, so give it a new Idempotency-Key, or a replay of the earlier key may return the original outcome instead of running your fix.
Sources
Related posts
More in Models
- Kling 4.0 Auto aspect ratio vs Sume's explicit aspect_ratio field
Kling 4.0 lists an Auto format next to 16:9, 1:1, 9:16 and 21:9. Sume's video ids take explicit ratios; see which ids list which values and how to read them.
- Kling 4.0's 3 to 30 seconds, split across Sume models by length
Kling 4.0 spans 3 to 30 s, and Flash 3 to 20 s. No single Sume id covers 3 to 30: Omni 3 to 10, Seedance 2.5 4 to 30, Wan 3.0 2 to 30. A lookup by seconds.
- Lyria 3.5 has no edit pass: iterate a music bed for $0.125 a take
Google says Lyria 3.5 is single-turn. Iterate a bed on Sume by changing one prompt axis per take; five takes cost $0.625 and a retry needs a key.
- Every Lyria 3.5 track carries SynthID: what brands should know
Google says all Lyria 3.5 output carries a SynthID audio watermark and blocks artist voices and copyrighted lyrics. Here is what that means for brand music.
Written by Sume