50-second AI video API: 30 + 20 seconds on Wan 3.0, cost by tier
A 50-second AI video on Sume takes two Wan 3.0 jobs, 30 s and 20 s: $3.13 at 480p, $6.25 at 720p, $12.50 at 1080p. Why Wan is the only model that can do it.

A 50-second AI video on Sume is two jobs: 30 seconds and 20 seconds. Wan 3.0 is the model to use, because it accepts any whole number from 2 to 30 seconds, so a 20-second second half is legal. The pair costs $3.13 at 480p, $6.25 at 720p and $12.50 at 1080p.
Models capped at 15 seconds would need four jobs for 50 seconds.
Which models can make the 20-second half?
Only the rows with a maximum above 15 seconds can produce a clip of 16 to 30 seconds in one job.
| Model id | Accepts 20 s | Longest |
|---|---|---|
| wan-3.0 | yes | 30 s |
| seedance-2.5 | yes | 30 s |
| seedance-2, seedance-2-fast, seedance-2-mini | no | 15 s |
| kling-3 | no | 15 s |
| grok-imagine-video-1.5 | no | 15 s |
| minimax-h3, minimax-h3-max | no | 15 s |
| gemini-omni-flash-1.1 | no | 10 s |
What does 30 + 20 cost on Wan 3.0?
Each job is billed at the list rate times 1.25 and rounded up to the cent. A 30-second job at 720p is $3.75 and a 20-second one is $2.50.
- 480p: $1.88 + $1.25 = $3.13
- 720p: $3.75 + $2.50 = $6.25
- 1080p: $7.50 + $5.00 = $12.50
How does it compare with four 15-second Kling jobs?
Kling 3 stops at 15 seconds, so 50 seconds means four jobs (for example 13, 13, 12 and 12 seconds), and Kling takes no reference images to hold a character across them. At $0.14 per second billed with audio off, 50 seconds of Kling is $7.00, close to Wan at 720p, but with three seams instead of one.
Fewer joins usually look better. Pick the model with the longest single clip that suits the look.
How do you request the pair?
Both halves are wan-3.0 jobs. The second takes a first frame cut from the end of the first.
import os, requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def job(key, secs, prompt, extra=None):
body = {"model": "wan-3.0", "prompt": prompt,
"duration": secs, "resolution": "720p", **(extra or {})}
r = requests.post("https://api.sume.com/v1/videos",
headers={**H, "Idempotency-Key": key}, json=body)
r.raise_for_status()
return r.json()["id"]
print(job("film-50-a", 30, "A cyclist crosses a city at dawn"))
print(job("film-50-b", 20, "The cyclist reaches a quiet harbour"))What should you check before you scale it?
Read usage.cost on the first completed 30-second job and compare with the table. If you plan many 50-second pieces, draft each half at 480p first.
Sources
Related posts
More in Pricing
- 60-second AI video API: two 30 s jobs or four 15 s jobs on Sume
A 60-second AI video on Sume is two 30 s jobs or four 15 s ones. Wan 3.0 costs $3.76, $7.50 or $15.00; four Kling 3 jobs cost $8.40. Priced and compared.
- Sizing generation_spend_cap_usd for a tool call from GPT-6.1 Sol
generation_spend_cap_usd has no default on Sume Agent Completions. Size it per run from metered API pricing and clamp it in your GPT-6.1 Sol tool handler.
- AI music cost per finished minute, compared
Lyria 3.5, ElevenLabs Music v2.5 and Suno Pro priced per minute or song, against Sume's flat $0.125 per accepted music generation.
- AI music for a 15-second bumper ad: per-minute vs flat pricing
A 15-second bumper needs 15 seconds of music. ElevenLabs at $0.15 a minute is about $0.04; Google lists $0.08 a song; Sume is a flat $0.125. When flat loses.
Written by Sume