Kling 3 on Sume rejects reference_*_urls: which model takes references
The kling-3 row supports text-to-video and start/end frames only. Send reference_image_urls and you get a 400; here is where references do work.

If you send reference_image_urls, reference_video_urls or reference_audio_urls to the kling-3 row on Sume, the request is refused: the row supports text-to-video or start/end frames only, with no reference_*_urls. The fix is not a retry. Either drop the references and use image_url plus end_image_url, or change the model to a row whose capabilities list references.
The error text reads model kling-3 supports text-to-video or start/end frames only (no reference_*_urls). and names the model, so it is easy to confuse with a typo in the field name. It is not.
What kling-3 does take
Sume's catalog describes kling-3 as Kling Video v3 Pro. Its capabilities say text-to-video, image-to-video and end frame are on, all three reference flags are off, native audio is on, resolutions are 720p and 1080p, duration is 4 to 15 seconds, and aspect ratios are 16:9, 9:16 and 1:1.
| Capability | Sume kling-3 row |
|---|---|
| Text-to-video | on |
| Image-to-video and end frame | on |
| Reference images, videos, audio | off: no reference_*_urls |
| Native audio | on |
| Resolutions | 720p, 1080p |
| Duration | 4 to 15 seconds |
Where references work instead
Rows with reference capabilities in the same catalog include wan-3.0, seedance-2.5, minimax-h3, minimax-h3-max and gemini-omni-flash-1.1. Their limits are not the same, so read each row before you move a prompt across.
wan-3.0: 10 images, 5 videos, 5 audio.minimax-h3andminimax-h3-max: 9 images, 3 videos, 3 audio, 12 in total.gemini-omni-flash-1.1: 10 images and 3 videos of 3 seconds or less, no audio references.
A quick guard in your client
Do the capability check before you call generate, so a batch does not burn calls on a known 400. This reads the catalog and keeps only rows that accept reference images.
import os
import requests
r = requests.get(
"https://api.sume.com/v1/video-router/models",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
timeout=30,
)
r.raise_for_status()
for row in r.json()["data"]:
caps = row.get("capabilities", {})
if caps.get("reference_images"):
print(row["id"])Check before you plan
The Video Router docs tell you to read capabilities from the catalog rather than assume one envelope, and the API reference has the full schema.
Sources
Related posts
More in Developers
- Koa receiver for Sume webhooks: verify with koa-bodyparser rawBody
koa-bodyparser keeps the raw string on ctx.request.rawBody. Check the sume-v1 HMAC against it, not against a re-stringified ctx.request.body, then answer 204.
- Korean karaoke captions: korean-ad and language ko on Sume
Burn Korean karaoke-style captions with style korean-ad and language ko on /v1/video-captions: one phrase at a time, the spoken word in a heavier weight.
- Workers KV jurisdictions: keep Sume job records in region
Cloudflare made Workers KV jurisdictions generally available on Oct 2, 2026. Here is how to key Sume job_id records into a region-scoped namespace.
- LangGraph custom image: keep SUME_API_KEY out of it
langgraph-cli 0.4.32 adds an --image-uri flag for self-hosted custom containers. Inject SUME_API_KEY as a runtime env var; never bake it into the image.
Written by Sume