sume/auto or a pinned video id after a vendor retires a model?

After the Sora removal, choose between model: sume/auto and a pinned catalog id. What each guarantees on Sume, with the 3-10 s, 720p, 8 s Auto defaults.

4 min readSume
All posts

Pin a catalog id when you must know which model made a clip, and use model: "sume/auto" when you only need a good clip and want Sume to keep choosing. Sume's docs say a replay of an sume/auto request gets the same price and the same route, and that the response reports sume/auto without naming the family.

OpenAI removed the Videos API and every Sora 2 id on 2026-09-24 (its deprecations page, read 2026-10-07), which is the clearest recent case for deciding this before a retirement, not during one.

What each choice guarantees

Both go through the same POST /v1/videos request. The difference is only the model field.

sume/auto vs a pinned id, from Sume's docs (read 2026-10-07)
Questionsume/autoPinned id such as gemini-omni-flash-1.1
Who chooses the modelSumeYou
Does the response name the familyNo, it reports sume/autoYes, the id you sent
Price on an idempotent replaySame price, same routeSame row price
Envelope3 to 10 s, 16:9 or 9:16, default 720p and 8 sWhatever the row's catalog entry says
Good forVariants where the model does not matterLook-matching, audits, reproducing a clip

A decision rule that fits on a card

If a client has ever asked which model made a shot, pin. If a clip is throwaway, such as a first-draft hook or a social variant you will test and discard, use Auto. If you do both, keep two config values and write the submitted model string on every job record.

  • Pinned ids give the reproducibility an approval trail needs.
  • Auto lets Sume change the family without a code change on your side.
  • Neither removes the need to read the response status and handle failed.

A Python submit that takes either

This sketch reads the model from the environment, so switching between sume/auto and a pinned id is a config change. It follows the documented submit and poll flow.

import os, time, requests

BASE = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

def main():
    body = {
        "model": os.environ.get("VIDEO_MODEL", "sume/auto"),
        "prompt": "A slow push-in on a ceramic mug on a wooden desk",
        "duration": 5,
    }
    job = requests.post(f"{BASE}/v1/videos", headers=H, json=body, timeout=30).json()
    while True:
        time.sleep(30)
        s = requests.get(job["polling_url"], headers=H, timeout=30).json()
        if s["status"] in ("completed", "failed", "cancelled"):
            print(s["status"], s.get("model"))
            break

main()

What to log

Log the requested model, the model in the poll response, the job id and usage.cost. With a pinned id the two model values match. With Auto the response says sume/auto, so keep your own record of why that choice was made.

Where Auto helps and where it hides things

Auto removes a decision, and that is both its use and its cost. You get a working clip without choosing a family, and a later change in what Sume routes to does not need your code to change. You lose the ability to say which family produced a given clip, since the docs state Sume never discloses it and tell callers not to infer it from the output.

That makes Auto a good fit for volume work where the clip is disposable and a poor fit for anything that is billed back to a client by model, or compared across a test run. If you are running a bake-off, pin every row; otherwise you are comparing Sume's routing, not the models.

  • Bake-offs and audits: pin.
  • Disposable variants and drafts: Auto is fine.
  • Look-matched campaigns: pin, and record the id with every job.

Cost and limits with Auto

Auto's create controls default to 720p and 8 seconds, with clips of 3 to 10 seconds in 16:9 or 9:16. If your brief needs 15 seconds or a square frame, Auto is outside its envelope and you must pin a row that supports it. The pinned rows list their limits in the catalog, and the price of each is its list rate times 1.25 per second.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume