Celery task for an AI video API: submit, poll, retry on Wan 3.0
Two Celery tasks for Sume's /v1/videos: one submits a Wan 3.0 job with an Idempotency-Key, one polls with self.retry(countdown) and stops on a terminal status.

A Celery worker should never sleep inside a task while an AI video renders. Split the work in two: a submit task that calls POST /v1/videos and returns in a second, and a poll task that reads the job once, then either finishes or calls self.retry(countdown=20) so the worker slot is free between reads. Sume answers the submit with 202 and a job id; the clip itself takes from about 30 seconds to several minutes, depending on the model, resolution and load (Sume video docs, read 2026-10-06).
The example below uses wan-3.0 at 720p, 9:16, 6 seconds. At Sume's Wan 3.0 rate of $0.125 per second at 720p that clip is $0.75, reserved at submit and shown as usage.cost on the finished job.
What do the two tasks look like?
Both use requests and the SUME_API_KEY environment variable. The Idempotency-Key is a value you own, such as a database row id, so a Celery redelivery of submit returns the original job instead of creating and billing a second one.
import os
import requests
from celery import Celery
app = Celery("clips", broker=os.environ["CELERY_BROKER_URL"])
API = "https://api.sume.com/v1/videos"
HEADERS = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
@app.task
def submit(prompt: str, key: str) -> str:
body = {"model": "wan-3.0", "prompt": prompt, "duration": 6,
"resolution": "720p", "aspect_ratio": "9:16"}
r = requests.post(API, json=body, timeout=30,
headers={**HEADERS, "Idempotency-Key": key})
r.raise_for_status()
job_id = r.json()["id"]
poll.apply_async((job_id,), countdown=20)
return job_id
@app.task(bind=True, max_retries=90)
def poll(self, job_id: str) -> str:
r = requests.get(f"{API}/{job_id}", headers=HEADERS, timeout=30)
r.raise_for_status()
job = r.json()
if job["status"] == "completed":
return f"{API}/{job_id}/content?index=0"
if job["status"] in ("failed", "cancelled"):
raise RuntimeError(job.get("error") or job["status"])
raise self.retry(countdown=20)Why does poll end with raise self.retry?
Celery's documentation says retry() raises a Retry exception, so no code after it runs, and that you can pass countdown to override the default delay, which is three minutes. With max_retries=90 and a 20 second countdown the task gives up after about 30 minutes, which is long for a 6 second Wan clip and generous for a 30 second 1080p one. When the limit is exceeded Celery re-raises the current exception, so a stuck job surfaces as a failed poll task rather than a silent loop (Celery tasks guide, read 2026-10-06).
A client-side timeout never cancels the render. Sume's jobs guide says the job keeps running and keeps billing; you only stopped waiting. Store the job id, and if you give up, cancel explicitly (only possible before generation starts) rather than submitting again.
What does one clip cost at each Wan resolution?
Wan 3.0 is billed per output second by resolution, list price times 1.25. The table is for the same 6 second clip the tasks submit.
| Resolution | List per second | Sume per second (list x 1.25) | 6-second clip |
|---|---|---|---|
| 480p | $0.05 | $0.0625 | $0.375 |
| 720p | $0.10 | $0.125 | $0.75 |
| 1080p | $0.20 | $0.25 | $1.50 |
What should the task do with the finished clip?
Download from the content endpoint with the same Authorization header: GET /v1/videos/{id}/content?index=0 redirects to the artifact. Sume's catalog also lists unsigned URLs on the poll response, but those need your key too, so a browser video tag cannot play them directly. Copy the file to your own storage in a third task and keep the Sume job id on the row so you can read /v1/jobs/{id}/events later.
If polling load matters, send callback_url on the submit and keep poll as the backup; the Django webhook view shows the receiving side.
Sources
Related posts
More in Developers
- Check a Shorts timeline body offline: a Python mirror of the rules
Catch Timeline 1.0 refusals you can see in the body (start at 0, 0.5 s coverage, transitions, even sizes) in 28 lines of Python before the plan call.
- Check an image request against the Sume catalog before you send it
Fetch GET /v1/images/models once, then reject a bad ratio, quality or reference count in Python before it reaches POST /v1/images and returns a 400.
- TTS then H3 Max lip sync: check the 5 to 14.8 s audio window
H3 Max lip sync takes 5 to 14.8 seconds of Sume-hosted audio and clips the rest. Measure a TTS line from its word timings and price it before you submit.
- CI smoke test for Ideogram 4.5 on Sume: one low-quality image
A bash and jq check that submits one 1K low-quality Ideogram 4.5 image, passes on 200 or 202, and explains 401, 402 and 429. List price is $0.03.
Written by Sume