Filter /v1/videos/models for a request in Python: 12 s, 1080p, 9:16

Read supported durations, resolutions, ratios and audio from Sume's model listing and keep only ids that can run your request. A 25-line script that runs.

5 min readSume
All posts

Hard-coded model tables go stale. Sume's GET /v1/videos/models returns, for each model, supported_durations, supported_resolutions, supported_aspect_ratios, generate_audio and pricing_skus, so a script can answer a request like 'a 12-second 1080p vertical clip with sound' from live data.

This is useful right now because new ids keep arriving. When Kling 4.0 or another model shows up in the listing, the same script will include it without a code change.

The script

It filters on exact membership in the supported lists and prints each match with its SKUs. Set SUME_API_KEY first.

import os
import requests

NEED = {"duration": 12, "resolution": "1080p", "aspect": "9:16", "audio": True}

def main() -> None:
    r = requests.get(
        "https://api.sume.com/v1/videos/models",
        headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
        timeout=30,
    )
    r.raise_for_status()
    for m in r.json()["data"]:
        fits = (
            NEED["duration"] in m["supported_durations"]
            and NEED["resolution"] in m["supported_resolutions"]
            and NEED["aspect"] in m["supported_aspect_ratios"]
            and (m["generate_audio"] or not NEED["audio"])
        )
        if fits:
            print(m["id"], m["pricing_skus"])

main()

Reading the output

pricing_skus has a different shape per model. Token-priced models (the Seedance family) return per-1000-video-tokens. Per-second models return per-video-second, with per-video-second-audio for Kling 3. Models that price by resolution return one key per tier, such as per-video-second-1080p. The values are the billable USD rate, already including Sume's 1.25 factor, so you can multiply by seconds. For tokens you need the pixel formula: width x height x seconds x 24 / 1024 tokens for Seedance.

Matching on membership also catches the traps a table hides. Omni Flash tops out at 10 seconds, so 12 does not appear in its list. Grok Imagine has no audio, so generate_audio is false and it drops out.

Next step

Add NEED["duration"] <= 10 style guards for reference inputs through supported_input_references, which lists image_url, video_url and audio_url per model. Then send the request with the chosen id. The same body works across ids, apart from the fields a model rejects.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume