Vendors swap GPUs; keep one video job shape

Luma said on Jul 23, 2026 it runs video-to-video inference on AMD and Tensorwave. Your client should not care: one Sume job shape covers every model.

4 min readSume
All posts

A Luma news item dated Jul 23, 2026 says Luma runs production inference on AMD and Tensorwave for video-to-video pipelines. Infrastructure moves like that are the vendor's business. Your integration should depend on a stable job shape instead, and on Sume every video model returns the same one.

What the news item says

The item is short, and only the claim below is used here. It says nothing about changes in price or output.

Luma news, Jul 23, 2026 (read 2026-10-03)
ItemPer the news item
Inference hardwareAMD and Tensorwave
WorkloadProduction video-to-video pipelines

What stays fixed on Sume

The Sume video docs describe one lifecycle for every catalog model: submit to POST /v1/videos, receive an id, polling_url and status, poll until completed, then download from unsigned_urls. The same job is also readable at GET /v1/jobs/{id}/status and /result. Statuses are a fixed set, and usage.cost is the billable amount.

Model ids are bare catalog ids with no provider prefix, and sume/auto never discloses which model family ran.

  • One submit route for all catalog models.
  • Same job lifecycle and status vocabulary.
  • Same status and result endpoints under /v1/jobs.
  • Idempotency-Key makes retries safe.

Write the client once

Keep the model id in configuration and the polling code model-agnostic. Then swapping a model, or a vendor moving its inference, is a config change with no client change. What does differ per model is capability: resolutions, durations and references, which you read from GET /v1/videos/models.

Check the capabilities per request, since a model that accepts 15 seconds will reject 30.

import os, time, requests

H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}

def render(model, prompt, **extra):
    r = requests.post("https://api.sume.com/v1/videos", headers=H, timeout=30,
                      json={"model": model, "prompt": prompt, **extra}).json()
    while True:
        s = requests.get(r["polling_url"], headers=H, timeout=30).json()
        if s["status"] in ("completed", "failed", "cancelled"):
            return s
        time.sleep(15)

# same code for any catalog id
print(render(os.environ.get("VIDEO_MODEL", "sume/auto"), "A mug on a desk")["status"])

Sources

Related posts

More in Developers

All Developers posts

Written by Sume