Kling 3 as your Sora replacement on Sume: four limits to check

Kling VIDEO 3.0 is 15 s with native audio per a video tracker. On Sume kling-3 takes 4-15 s at 720p or 1080p and no reference URLs. Four checks before a port.

5 min readSume
All posts

Sume lists Kling 3 as kling-3, and it takes 4-15 seconds at 720p or 1080p. The Video Router doc says it accepts no reference_*_urls. A video tracker records Kling VIDEO 3.0 as released on 2026-02-05, 15 seconds long, with native audio (Magic Hour tracker, read 2026-10-06). The same tracker records that the OpenAI Sora API ended on 2026-09-24.

So Kling 3 can stand in for a Sora call that made 4 to 15 second clips from text or a first frame. It cannot stand in for a call that passed reference images to hold a subject. Check four things before you point code at it.

The four checks

kling-3 on Sume, Video Router doc, read 2026-10-06
CheckWhat the doc saysWhat to do in your code
Duration4-15 sClamp or split anything over 15 s; under 4 s needs another model
Resolution720p and 1080pDrop 480p drafts to another model
ReferencesNo reference_*_urlsUse frame_images for a first frame, not references for a subject
Aspect ratioNot stated in the router tableRead supported_aspect_ratios before you ask for 9:16

Run the four checks in code

The function reads the catalog row and returns every problem it finds for a given request, so a bad request never reaches the paid endpoint.

import os
import requests


def problems(model: str, req: dict) -> list[str]:
    r = requests.get(
        "https://api.sume.com/v1/videos/models",
        headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
        timeout=30,
    )
    r.raise_for_status()
    m = next(x for x in r.json()["data"] if x["id"] == model)
    out = []
    if req.get("duration") not in (m.get("supported_durations") or []):
        out.append("duration not supported")
    if req.get("resolution") not in (m.get("supported_resolutions") or []):
        out.append("resolution not supported")
    if req.get("aspect_ratio") not in (m.get("supported_aspect_ratios") or []):
        out.append("aspect ratio not supported")
    if req.get("input_references") and not m.get("supported_input_references"):
        out.append("input_references not supported")
    return out


print(problems("kling-3", {"duration": 8, "resolution": "720p", "aspect_ratio": "9:16"}))

Where Kling 3 fits and where it does not

  • Fits: a single 4 to 15 second shot from text, or from a first frame, at 720p or 1080p.
  • Does not fit: shots shorter than 4 seconds, which Omni (3-10 s) or Wan (2-30 s) take; shots longer than 15 seconds, which Seedance 2.5 and Wan take.
  • Does not fit: reference-image workflows. Use a model whose supported_input_references lists image, such as Seedance 2.0 or Omni.
  • Pricing: the router table in the clone does not give a per-second list price for kling-3, so read pricing_skus from the catalog rather than quoting a number.

Run one prompt through it and two other models and compare, as in the earlier migration table post that maps the usual names to Sume ids (migration table). If your first frame is a photo, the Kling 3 first-frame request post has the body to start from.

Keep the model id in config, not in code. A catalog can add, rename or retire an id, and a one-line config change is cheaper than a redeploy.

Making the swap reversible

Point code at a new model behind a switch, not as a find-and-replace. Put the model id, and the limits you rely on, in one config object. The function that builds the request reads the config and runs the four checks. If Kling 3 turns out to be wrong for one job type, you change one value for that type and leave the others on their model.

Run a small share of traffic through it first. A ten percent canary, with the five-number scorecard on each side, tells you in a day or two whether the swap holds up on your own prompts, which no tracker page can do for you.

If a check fails, the right response is usually to route that job type to another listed model rather than to bend the request. Clamping a 20 second ask to 15 seconds silently changes the product, while a clear refusal in your own code tells the caller what to change.

  • Keep the model id in config, per job type, not per deployment.
  • Run the four checks at request time and return the problems to the caller instead of sending a bad request.
  • Log the model id on every job, so a later complaint can be traced to a model.
  • Read the price from the catalog's pricing_skus and not from a blog, since the router table does not list a per-second price for this model.

Sources

Related posts

More in Models

All Models posts

Written by Sume