Poll an H3 Max video job in Python: backoff, terminal, result

Submit minimax-h3-max to the Video Router, poll GET /v1/jobs/:id/status with backoff until terminal, then fetch the result. A runnable 25-line Python script.

5 min readSume
All posts

Submit a minimax-h3-max job to POST /v1/video-router/generate with mode: "async", keep the job id, and poll GET /v1/jobs/:id/status until the response says it is terminal. Then read GET /v1/jobs/:id/result.

The H3 Max row takes 5 to 15 seconds at 480p, 768p or 1080p, and a 10-second 768p clip is $1.00 on Sume, so the script below logs the quoted price from the submit response before it starts to wait.

The script

This uses requests and an Idempotency-Key, so running it twice with the same key does not create a second paid job. Set SUME_API_KEY first.

import os, time, requests

API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

body = {
    "model": "minimax-h3-max",
    "prompt": "A ceramic mug on a desk, slow push-in, soft morning light",
    "resolution": "768p",
    "duration": 10,
    "mode": "async",
}
r = requests.post(f"{API}/v1/video-router/generate", json=body, timeout=30,
                  headers={**H, "Idempotency-Key": "h3max-poll-demo-001"})
r.raise_for_status()
sub = r.json()
job_id = sub["job"]["id"]
print("quoted micros:", sub.get("usage", {}).get("billable_amount_usd_micros"))

delay = 3.0
while True:
    s = requests.get(f"{API}/v1/jobs/{job_id}/status", headers=H, timeout=30).json()
    s = s.get("job", s)
    if s.get("terminal"):
        break
    time.sleep(delay)
    delay = min(delay * 1.5, 30.0)

print("status:", s.get("status"))
if s.get("status") == "completed":
    res = requests.get(f"{API}/v1/jobs/{job_id}/result", headers=H, timeout=30)
    print(res.json())

What the loop relies on

The Sume jobs docs say that in async mode you poll status_url until terminal is true, then fetch result_url when result_ready is true. Terminal statuses are completed, failed and canceled; non-terminal statuses are queued and processing. The result is a conflict response while the job is still running, not an empty body, so do not fetch it early.

The backoff starts at 3 seconds and grows by half each time up to 30 seconds. The docs do not publish a render-time figure for H3 Max, so the script does not guess one; it waits until the job reaches a terminal state.

What can go wrong

If the submit call times out, retry with the same Idempotency-Key. Sume's docs say a key reused with a different body is a conflict, so keep the body identical on a retry. If a job fails, the reservation is refunded; check status for failed and read the error before you resubmit.

The sync mode waits at most 30 seconds on the submit call. A 10 to 15 second H3 Max clip may outlast that, and the docs say to poll and not resubmit when a sync call comes back non-terminal.

  • Do not poll in a tight loop; use the backoff.
  • Do not resubmit when the status is queued or processing.
  • Log the job id and the quoted micros for every call.

Use a webhook for batches

If you queue dozens of clips, polling each one adds up. Send mode: "webhook" with a public HTTPS webhook_url and Sume sends signed job.completed, job.failed and job.canceled events. The docs advise keeping a poll as a backup. Verify the signature before you trust the payload, and refuse to run with an empty secret.

Reading the price from the submit response

The standard submit envelope carries job.id, status_url, result_url and usage.billable_amount_usd_micros. A micro is one millionth of a dollar, so $1.00 is 1,000,000 micros. For a 10-second 768p H3 Max clip the quote should be 1,000,000, because 10 seconds x $0.08 list x 1.25 is $1.00. If you see a different number, check the resolution and duration you sent against the catalog row.

Keep the quote next to the job id in your own log. When the job completes, the captured amount should equal the quote; on failure it is refunded.

Choosing the poll interval

Sume does not publish a latency promise for H3 Max in the docs this post draws on, so the interval is a policy choice. Starting at 3 seconds and stretching to 30 keeps the request count per job to a few dozen at most, even for a 15-second clip at 1080p. If you queue many clips, spread the polls and prefer the webhook for the terminal signal.

Canceling and recovering

If you submit by mistake, you can cancel with POST /v1/jobs/:id/cancel, but only before generation starts. After generation begins the API answers 409 job_generation_already_started and the job runs to completion, so cancel quickly or accept the charge. A cancel on an already canceled job returns the same canceled job.

If your process dies after the submit, you do not lose the job. List recent jobs with GET /v1/jobs, filter by status, and read GET /v1/jobs/:id/events for the timeline: job.created, job.queued, job.started, generation.submitted and then job.completed or job.failed. That is also the quickest way to see whether a retry with the same Idempotency-Key returned the earlier job.

Status fields to read

The status endpoint reports one of five states. Stop on the terminal ones and fetch the result only when it is ready.

Job statuses for an H3 Max poll loop (Sume Jobs docs read 2026-10-05)
StatusTerminalAction
queuedNoWait and poll again
processingNoWait and poll again
completedYesFetch /v1/jobs/:id/result
failedYesRead the error; the charge is refunded
canceledYesStop; nothing to fetch

Sources

Related posts

More in Developers

All Developers posts

Written by Sume