60-second AI video API: two 30 s jobs or four 15 s jobs on Sume
A 60-second AI video on Sume is two 30 s jobs or four 15 s ones. Wan 3.0 costs $3.76, $7.50 or $15.00; four Kling 3 jobs cost $8.40. Priced and compared.

A 60-second AI video on Sume takes at least two jobs, because no model goes beyond 30 seconds. The cheapest route is two Wan 3.0 jobs of 30 seconds: $3.76 at 480p, $7.50 at 720p and $15.00 at 1080p.
The alternative is four 15-second jobs on a 15-second model, which gives you more cut points and, on Kling 3, costs $8.40 with audio off.
What does each way of building 60 seconds cost?
Sume bills the provider list times 1.25 and rounds each job up to the cent, so the totals below are job totals added together.
| Route | Jobs | Tier | Total |
|---|---|---|---|
| wan-3.0, 30 s each | 2 | 480p | $3.76 |
| wan-3.0, 30 s each | 2 | 720p | $7.50 |
| wan-3.0, 30 s each | 2 | 1080p | $15.00 |
| kling-3, 15 s each, audio off | 4 | 720p or 1080p | $8.40 |
| kling-3, 15 s each, audio on | 4 | 720p or 1080p | $12.60 |
| seedance-2.5, 30 s each | 2 | 720p | $34.68 |
Is two big jobs better than four small ones?
One seam is easier to hide than three. Two 30-second jobs give a single cut at the midpoint and keep both halves on a model that takes reference images. Four 15-second jobs suit a piece with natural scene changes every 15 seconds, such as a four-step explainer.
Kling 3 takes no reference images, so four Kling jobs rely on the first frame to carry continuity. Wan, Seedance and the MiniMax rows accept input_references for the same subject across jobs.
How do you keep the look consistent?
Pass the same reference images to every job, or chain the jobs by starting each from the final frame of the one before. The request below uses the first approach with Wan 3.0.
- Same model, resolution and aspect ratio on every job.
- Reuse the same
input_referencesentries across jobs. - Use an
Idempotency-Keyper job so a retry never bills twice.
What does the request look like?
Two jobs of 30 seconds each, keys read from the environment.
import os, requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
ref = {"type": "image_url", "image_url": {"url": os.environ["REF_URL"]}}
scenes = ["Opening: she enters the workshop",
"Closing: she finishes the piece and smiles"]
for i, scene in enumerate(scenes):
r = requests.post(
"https://api.sume.com/v1/videos",
headers={**H, "Idempotency-Key": f"film-60-{i}"},
json={"model": "wan-3.0", "prompt": scene, "duration": 30,
"resolution": "720p", "input_references": [ref]},
)
r.raise_for_status()
print(i, r.json()["id"])What does it cost to test first?
A 5-second 480p draft is $0.32 on Wan. Two drafts cost $0.64, well under 10 percent of the $7.50 final.
Sources
Related posts
More in Pricing
- Sizing generation_spend_cap_usd for a tool call from GPT-6.1 Sol
generation_spend_cap_usd has no default on Sume Agent Completions. Size it per run from metered API pricing and clamp it in your GPT-6.1 Sol tool handler.
- AI music cost per finished minute, compared
Lyria 3.5, ElevenLabs Music v2.5 and Suno Pro priced per minute or song, against Sume's flat $0.125 per accepted music generation.
- AI music for a 15-second bumper ad: per-minute vs flat pricing
A 15-second bumper needs 15 seconds of music. ElevenLabs at $0.15 a minute is about $0.04; Google lists $0.08 a song; Sume is a flat $0.125. When flat loses.
- What a three-minute AI song costs: ElevenLabs, Lyria and Sume
At list prices a three-minute track is $0.45 on ElevenLabs Music API ($0.15 per minute), $0.08 on Lyria 3.5 and $0.125 on Sume's Music Router.
Written by Sume