Python: three Wan 3.0 hooks from one reference image, with costs
A Python script for Sume's /v1/videos: submit three Wan 3.0 hook prompts with one reference image at 480p, poll each job, and print the usage cost.

To test three hook ideas against one reference image, submit three wan-3.0 jobs to POST /v1/videos with the same input_references entry and different prompts, poll each polling_url until it completes, and print usage.cost from the poll response. The request fields and the poll fields in the script below are the ones in Sume's Video generation docs. Wan 3.0 is the right model for a cheap first pass because the docs list its range as 2 to 30 seconds, so a 2 or 3 second hook is a valid request; the catalog also lists a 480p tier for drafts.
The script
It needs requests and the environment variable SUME_API_KEY. The reference URL is a placeholder: replace it with an image you host. It submits three jobs first, so they run in parallel, then polls them. Poll interval follows the docs' example of 30 seconds.
import os, time, requests
KEY = os.environ["SUME_API_KEY"]
H = {"Authorization": f"Bearer {KEY}", "Content-Type": "application/json"}
URL = "https://api.sume.com/v1/videos"
REF = {"type": "image_url", "image_url": {"url": "https://example.com/product.png"}}
HOOKS = {
"question": "A hand holds the product up to the camera, quick push-in, bright kitchen.",
"reveal": "The product slides out of a box in slow motion, soft top light.",
"before-after": "Split scene: messy desk, then the same desk tidy with the product.",
}
jobs = {}
for name, prompt in HOOKS.items():
r = requests.post(URL, headers=H, json={
"model": "wan-3.0", "prompt": prompt, "duration": 3,
"resolution": "480p", "aspect_ratio": "9:16", "input_references": [REF]})
r.raise_for_status()
jobs[name] = r.json()["polling_url"]
while jobs:
time.sleep(30)
for name, url in list(jobs.items()):
s = requests.get(url, headers=H).json()
if s["status"] == "completed":
print(name, "done, cost", s.get("usage", {}).get("cost"), s["unsigned_urls"][0])
del jobs[name]
elif s["status"] == "failed":
print(name, "failed:", s.get("error", "unknown"))
del jobs[name]What the script assumes
It follows the documented flow: submit, poll the polling_url, and read unsigned_urls[0] when the status is completed. The status values completed and failed come from the docs' own example. The docs show a single reference as an object with type: "image_url" and a nested image_url.url, which is what REF is. A model accepts a reference type only if supported_input_references lists it, and the docs say Wan 3.0 accepts image, video and audio references.
| Field | Value in the script | Source |
|---|---|---|
| model | wan-3.0 | Video generation docs, request parameters |
| duration | 3 (seconds, integer) | wan-3.0 accepts 2 to 30 s |
| resolution | 480p | Catalog tier, checked in supported_resolutions |
| aspect_ratio | 9:16 | Listed in the docs for vertical output |
| input_references | One image_url object | Reference-to-video section |
| usage.cost | Printed per job | Poll response example |
Make it safer
Two things are missing on purpose. The docs' example does not send an idempotency header on /v1/videos, so a retry after a network error could submit the job twice; if you add retry logic, check the docs for the header first. And the loop has no timeout: add a deadline so a stuck job cannot hold the script forever. Then pick the winner by eye, and regenerate only that hook at 720p.
Why three variants and one image. A hook test is only fair if the only thing that changes is the idea. Holding the reference image and the length fixed means a difference in the clips comes from the prompt, and a difference in cost comes from nothing at all, since the price follows seconds and resolution. If one variant costs more than another, something other than the prompt changed, and the printed usage.cost is the first place to look.
Reading the output. The script prints the name, the cost and the first content URL for each finished job. Open the three files side by side, and judge the first second of each, since that is what a hook is. Write down which wording of the prompt produced the better opening, and reuse that wording in the next round. Keep the three prompts in a file next to the clips so you can tell later which prompt made which file, and delete the clips you will not use so you do not pay to store them.
Scale it up. To test ten hooks, change only the HOOKS dictionary. The loop already submits all jobs before it polls, so the wall-clock time is about that of the slowest job. The Sume docs describe queue-first admission with workspace concurrency limits, so if you submit more jobs than your workspace can run at once, the extra jobs wait in a queue; check your limits before a large batch.
Sources
Related posts
More in Developers
- Python TTS cost calculator: MAI-Voice-2.1, Flash and Sume per job
A short Python function prices any script on MAI-Voice-2.1 ($22/M), Flash ($15/M) and Sume (list x 1.25, rounded up per job); 210 vs 211 characters shown.
- Test a video poll loop with unittest and a local server, no spend
A 30-line stdlib file tests a Sume video poll loop against a fake server: pending, in_progress, completed in order, and a failed job that stops with no sleep.
- Python urllib only: a Wan 3.0 test job for $0.125
No requests, no SDK: 27 lines of Python urllib submit a 2-second 480p Wan 3.0 job, poll it and save the MP4. The job costs $0.125, the cheapest Wan clip.
- Queue a 30 s H3 Max Recast job: the reserve at 768p and 1080p
H3 Max Recast swaps people in a 5 to 30 s video. Price arithmetic for a 30 s source at 768p and 1080p, plus the Python to submit it and poll it.
Written by Sume