Hosted video instead of a GPU: a Python submit-poll-download script
Replace a local H3 or Wan pipeline with three HTTP calls. A Python script that submits a minimax-h3 job to Sume, polls it and saves the MP4, with job states.

If you only need the clip, the hosted version of an open-weights pipeline is three HTTP calls: POST /v1/videos to submit, GET on the returned polling URL until the status is completed, then download from unsigned_urls[0]. The script below does that for minimax-h3 at 768p for five seconds, which Sume prices at $0.375 (5 x $0.06 list x 1.25). It uses the requests package and your SUME_API_KEY. Fields are from the Sume video docs as of 2026-10-08.
The script
It sends an Idempotency-Key so a retried submit returns the original job, polls every 30 seconds as the docs suggest, and stops on failed or cancelled.
import os, time, requests
KEY = os.environ["SUME_API_KEY"]
H = {"Authorization": f"Bearer {KEY}",
"Idempotency-Key": "h3-clip-001"}
body = {"model": "minimax-h3", "resolution": "768p",
"duration": 5, "aspect_ratio": "16:9",
"prompt": "A paper boat drifts down a rain gutter"}
r = requests.post("https://api.sume.com/v1/videos",
headers=H, json=body, timeout=60)
r.raise_for_status()
job = r.json()
print(job["id"], job["status"])
while True:
time.sleep(30)
s = requests.get(job["polling_url"], headers=H, timeout=60).json()
print(s["status"])
if s["status"] == "completed":
break
if s["status"] in ("failed", "cancelled"):
raise SystemExit(s.get("error", s["status"]))
v = requests.get(s["unsigned_urls"][0], headers=H, timeout=300)
open("clip.mp4", "wb").write(v.content)
print("saved clip.mp4")Job states and what the script does
The docs list five statuses.
| Status | Meaning | Script action |
|---|---|---|
| pending | Submitted and in the queue | Keep polling |
| in_progress | Generation is running | Keep polling |
| completed | The video can be downloaded | Download unsigned_urls[0] |
| failed | Generation failed; see the error field | Exit with the error |
| cancelled | The job is cancelled and did not complete | Exit |
Limits to know before you scale it
Sume says minimax-h3 accepts 5 to 15 seconds at native 480p or 768p. A seed field is rejected, since no v1 model accepts one, and size returns a 400, so use resolution plus aspect_ratio. A non-empty provider.options is also rejected. The docs say video generation usually takes 30 seconds to several minutes depending on model and parameters, so the loop above is built for waiting, not for an interactive request.
For many jobs, send a callback_url (HTTPS) and let Sume POST a signed webhook when a job reaches a terminal state instead of polling. The signature is in the x-sume-webhook-signature header with a timestamp header; refuse to verify with an empty secret.
Swapping the model
Change model and resolution and the same script works for another listed id, provided the duration is in range.
wan-3.0: 2 to 30 seconds, 720p at $0.125 per second.minimax-h3-max: 5 to 15 seconds, 768p at $0.10 per second.- Not listed on Sume: Wan 2.2, HunyuanVideo and LTX models.
What to change for production
Store the job id before you start polling, so a crashed process can resume with GET /v1/videos/{jobId}. Use a new Idempotency-Key per distinct clip, because a replay with the same key returns the original job rather than starting a new one. Catch requests errors around the poll loop, and cap the number of polls so a stuck job cannot hold a worker forever.
Why not a GPU for this
A local H3 or Wan 2.2 pipeline needs a download of tens of gigabytes, a CUDA machine and a license check; the script above needs none of those. It also bills only for jobs you submit: at $0.075 per second on minimax-h3 at 768p, ten clips of five seconds cost $3.75 (10 x 5 x $0.06 x 1.25). Sume documents that the price is reserved at submit, so a loop that submits in error is visible in your balance.
Sources
Related posts
More in Developers
- How many Sume jobs can one key poll at once? Reads per minute by plan
Polling every 2 seconds costs 30 reads per job per minute. Divide your plan's read budget by 30 to size the poller: 160 jobs on Free, 1600 on Scale.
- How many minutes of speech fit in one 20,000-character TTS request?
Sume TTS caps a request at 20,000 characters. At 800 to 1,200 characters per minute that is roughly 17 to 25 minutes of audio and costs 95 cents.
- Hy Image 3.5 Preview returns base64 PNG; Sume returns a URL
Moving image code from a base64 PNG response (Hy Image 3.5 Preview on OpenRouter) to Sume means downloading data[].url. A 12-line Python version.
- One idempotency key per prompt row: re-run a 1,000-clip library safely
Derive the Idempotency-Key from every field you send, so a crashed batch can restart without paying twice and a changed row never collides. Python, 17 lines.
Written by Sume