Python: three Wan 3.0 hooks from one reference image, with costs

A Python script for Sume's /v1/videos: submit three Wan 3.0 hook prompts with one reference image at 480p, poll each job, and print the usage cost.

5 min readSume
All posts

To test three hook ideas against one reference image, submit three wan-3.0 jobs to POST /v1/videos with the same input_references entry and different prompts, poll each polling_url until it completes, and print usage.cost from the poll response. The request fields and the poll fields in the script below are the ones in Sume's Video generation docs. Wan 3.0 is the right model for a cheap first pass because the docs list its range as 2 to 30 seconds, so a 2 or 3 second hook is a valid request; the catalog also lists a 480p tier for drafts.

The script

It needs requests and the environment variable SUME_API_KEY. The reference URL is a placeholder: replace it with an image you host. It submits three jobs first, so they run in parallel, then polls them. Poll interval follows the docs' example of 30 seconds.

import os, time, requests

KEY = os.environ["SUME_API_KEY"]
H = {"Authorization": f"Bearer {KEY}", "Content-Type": "application/json"}
URL = "https://api.sume.com/v1/videos"
REF = {"type": "image_url", "image_url": {"url": "https://example.com/product.png"}}
HOOKS = {
    "question": "A hand holds the product up to the camera, quick push-in, bright kitchen.",
    "reveal": "The product slides out of a box in slow motion, soft top light.",
    "before-after": "Split scene: messy desk, then the same desk tidy with the product.",
}

jobs = {}
for name, prompt in HOOKS.items():
    r = requests.post(URL, headers=H, json={
        "model": "wan-3.0", "prompt": prompt, "duration": 3,
        "resolution": "480p", "aspect_ratio": "9:16", "input_references": [REF]})
    r.raise_for_status()
    jobs[name] = r.json()["polling_url"]

while jobs:
    time.sleep(30)
    for name, url in list(jobs.items()):
        s = requests.get(url, headers=H).json()
        if s["status"] == "completed":
            print(name, "done, cost", s.get("usage", {}).get("cost"), s["unsigned_urls"][0])
            del jobs[name]
        elif s["status"] == "failed":
            print(name, "failed:", s.get("error", "unknown"))
            del jobs[name]

What the script assumes

It follows the documented flow: submit, poll the polling_url, and read unsigned_urls[0] when the status is completed. The status values completed and failed come from the docs' own example. The docs show a single reference as an object with type: "image_url" and a nested image_url.url, which is what REF is. A model accepts a reference type only if supported_input_references lists it, and the docs say Wan 3.0 accepts image, video and audio references.

Script fields and where they come from (read 2026-10-05)
FieldValue in the scriptSource
modelwan-3.0Video generation docs, request parameters
duration3 (seconds, integer)wan-3.0 accepts 2 to 30 s
resolution480pCatalog tier, checked in supported_resolutions
aspect_ratio9:16Listed in the docs for vertical output
input_referencesOne image_url objectReference-to-video section
usage.costPrinted per jobPoll response example

Make it safer

Two things are missing on purpose. The docs' example does not send an idempotency header on /v1/videos, so a retry after a network error could submit the job twice; if you add retry logic, check the docs for the header first. And the loop has no timeout: add a deadline so a stuck job cannot hold the script forever. Then pick the winner by eye, and regenerate only that hook at 720p.

Why three variants and one image. A hook test is only fair if the only thing that changes is the idea. Holding the reference image and the length fixed means a difference in the clips comes from the prompt, and a difference in cost comes from nothing at all, since the price follows seconds and resolution. If one variant costs more than another, something other than the prompt changed, and the printed usage.cost is the first place to look.

Reading the output. The script prints the name, the cost and the first content URL for each finished job. Open the three files side by side, and judge the first second of each, since that is what a hook is. Write down which wording of the prompt produced the better opening, and reuse that wording in the next round. Keep the three prompts in a file next to the clips so you can tell later which prompt made which file, and delete the clips you will not use so you do not pay to store them.

Scale it up. To test ten hooks, change only the HOOKS dictionary. The loop already submits all jobs before it polls, so the wall-clock time is about that of the slowest job. The Sume docs describe queue-first admission with workspace concurrency limits, so if you submit more jobs than your workspace can run at once, the extra jobs wait in a queue; check your limits before a large batch.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume