Split a 15-reference storyboard into two Wan 3.0 jobs on Sume
A 15-image brief does not fit one Sume Wan 3.0 job, which takes 10 references. Split it into two jobs in Python; two 8 s 720p jobs cost $2.00.

If your brief has 15 reference images, one Sume Wan 3.0 job cannot take all of them: the catalog limit is 10 reference images per job. Split the list into two jobs of 8 and 7 images, keep the character images in both, and join the two clips afterwards. Two 8-second jobs at 720p cost $1.00 each, so $2.00 for the pair. The short Python script below does the split and submits both with idempotency keys.
Why 8 and 7, not 10 and 5
An even split leaves room to repeat shared images. If the lead character appears in both halves, you want that image in both jobs, which is a slot used twice. With 15 distinct images, 8 plus 7 leaves two free slots in the first job and three in the second for the repeated character and product images, up to the cap of 10.
Wan 3.0 also accepts reference videos and audio, but this script uses images only, as the Vidu limit under discussion is for images.
| Item | Value | Source |
|---|---|---|
| Reference images per job | up to 10 | Sume Video Router catalog |
| Duration | 2 to 30 seconds | Sume video docs |
| Rate at 720p | $0.125 a second | Provider list $0.10 x 1.25 |
| One 8 s job at 720p | $1.00 | 8 x $0.125 |
| Two jobs | $2.00 | 2 x $1.00 |
The script
It takes a shared list (character images that go into every job) and a list of shot images, packs them into jobs of at most 10, and submits each. Set SUME_API_KEY first.
import os
import requests
KEY = os.environ["SUME_API_KEY"]
H = {"Authorization": f"Bearer {KEY}"}
CAP = 10
shared = ["https://example.com/lead.png", "https://example.com/product.png"]
shots = [f"https://example.com/shot-{i}.png" for i in range(13)]
room = CAP - len(shared)
batches = [shots[i:i + room] for i in range(0, len(shots), room)]
for n, batch in enumerate(batches):
urls = shared + batch
body = {
"model": "wan-3.0",
"prompt": f"Part {n + 1}: the founders move through the studio",
"resolution": "720p",
"duration": 8,
"input_references": [
{"type": "image_url", "image_url": {"url": u}} for u in urls
],
}
r = requests.post(
"https://api.sume.com/v1/videos",
headers={**H, "Idempotency-Key": f"split-demo-{n}"},
json=body, timeout=60,
)
print(n, len(urls), r.status_code, r.json().get("id"))After the split
Steps:
- Poll both jobs with
GET /v1/videos/{id}until each iscompleted. - Download each from
unsigned_urls[0]. - Cut the two clips together in your editor, or use Sume's timeline compose.
- Match the two prompts so lighting and wardrobe words agree across the cut.
Writing the two prompts
The prompts carry continuity, so write them as a pair. Start both with the same sentence about setting, light and wardrobe, then add the action that is unique to each half. If the lead wears a green jacket in part one, say so in part two as well, because the second job knows nothing about the first.
Order the shared images the same way in both jobs. Put the lead character first, the product second, and the shot-specific images after them. A consistent order makes it easier to compare outputs and, on rows that document prompt tags, keeps the indexes stable.
Cost check before you run
Use dry_run equivalents where they exist. On the plain REST route, a quick estimate is enough: duration times the per-second rate, rounded up to the cent per job. At 1080p the same pair would be 2 x 8 x $0.25, or $4.00, so decide the resolution before you split. Draft at 480p, which is $0.50 a job after rounding, and re-render only the approved half at full size.
What Sume does not do
Sume does not stitch the two outputs into one continuous shot by itself, and the two jobs will not share a latent state, so small differences in the characters can show at the cut. Hide the join behind a camera move or a cutaway. It also does not accept a Vidu model name, so this approach is for readers who want the closest Sume equivalent of a many-reference brief.
Sources
Related posts
More in Developers
- Streamlit Sora demo: switch to Sume and st.video, no double billing
Streamlit reruns your script on every click. Derive the Sume Idempotency-Key from the prompt so a rerun returns the first job, then show the clip with st.video.
- Sume 429: retry-after header first, then the body field, then backoff
How long to wait after a Sume 429? Use the retry-after header, then error.retry_after_seconds, then jittered backoff. A Python stdlib function with checks.
- How many images can I attach to a Sume agent run? 30, up to 500 MB
An Agent Completion takes at most 30 images, each up to 30 MB, 500 MB total. So 30 images of 30 MB cannot all fit; 16 can. Error codes inside.
- Which Sume key scope reads or rotates the webhook signing secret?
Reading the webhook signing secret needs account:read; rotating needs account:write. Scopes are fixed at key creation, so use a separate key.
Written by Sume