Vendors swap GPUs; keep one video job shape
Luma said on Jul 23, 2026 it runs video-to-video inference on AMD and Tensorwave. Your client should not care: one Sume job shape covers every model.

A Luma news item dated Jul 23, 2026 says Luma runs production inference on AMD and Tensorwave for video-to-video pipelines. Infrastructure moves like that are the vendor's business. Your integration should depend on a stable job shape instead, and on Sume every video model returns the same one.
What the news item says
The item is short, and only the claim below is used here. It says nothing about changes in price or output.
| Item | Per the news item |
|---|---|
| Inference hardware | AMD and Tensorwave |
| Workload | Production video-to-video pipelines |
What stays fixed on Sume
The Sume video docs describe one lifecycle for every catalog model: submit to POST /v1/videos, receive an id, polling_url and status, poll until completed, then download from unsigned_urls. The same job is also readable at GET /v1/jobs/{id}/status and /result. Statuses are a fixed set, and usage.cost is the billable amount.
Model ids are bare catalog ids with no provider prefix, and sume/auto never discloses which model family ran.
- One submit route for all catalog models.
- Same job lifecycle and status vocabulary.
- Same status and result endpoints under /v1/jobs.
- Idempotency-Key makes retries safe.
Write the client once
Keep the model id in configuration and the polling code model-agnostic. Then swapping a model, or a vendor moving its inference, is a config change with no client change. What does differ per model is capability: resolutions, durations and references, which you read from GET /v1/videos/models.
Check the capabilities per request, since a model that accepts 15 seconds will reject 30.
import os, time, requests
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
def render(model, prompt, **extra):
r = requests.post("https://api.sume.com/v1/videos", headers=H, timeout=30,
json={"model": model, "prompt": prompt, **extra}).json()
while True:
s = requests.get(r["polling_url"], headers=H, timeout=30).json()
if s["status"] in ("completed", "failed", "cancelled"):
return s
time.sleep(15)
# same code for any catalog id
print(render(os.environ.get("VIDEO_MODEL", "sume/auto"), "A mug on a desk")["status"])Sources
Related posts
More in Developers
- Animate a still by API: first-frame jobs on Sume
fal lists FLUX 3 as animating one still into video. On Sume you send the still as a first frame to a video model; this Python script submits and polls.
- Ask for the aspect ratio first: an MCP input_required round trip
MCP multi round-trip requests let a tool answer input_required to ask for an aspect ratio or spend approval before a render, with state in requestState.
- Attach a terminal to a run: ant sessions connect vs Sume jobs
The ant CLI can attach to a Managed Agents session. For Sume generation jobs, use sume jobs watch and MCP jobs_wait instead, and never resubmit a paid job.
- Audit logs without file names: what to log for Sume jobs
Claude's Compliance API Activity Feed stopped returning file names. For Sume jobs, log request ids and job ids, never media URLs or transcripts.
Written by Sume