31-second AI video API: why Sume needs two jobs past 30 seconds
No Sume video model takes duration 31. Wan 3.0 and Seedance 2.5 stop at 30 seconds; here is the two-job route, what it costs and how to join the clips.

A single 31-second video is not possible on Sume. The longest duration any video model accepts is 30 seconds, on wan-3.0 and seedance-2.5, so a duration: 31 request is refused with a 400. To get 31 seconds you make two jobs and join them.
The cheapest split is a 30-second Wan 3.0 job plus a short second clip, and the cost of the pair is just the sum of the two job prices.
Which models can go past 15 seconds at all?
Only two catalog models run beyond 15 seconds, and one more specialised row reaches 30. Everything else tops out at 10 or 15.
| Model id | Longest | Shortest |
|---|---|---|
| wan-3.0 | 30 s | 2 s |
| seedance-2.5 | 30 s | 4 s |
| h3-max-recast | 30 s | 5 s |
| seedance-2, seedance-2-fast, seedance-2-mini | 15 s | 4 s |
| kling-3, grok-imagine-video-1.5 | 15 s | 4 s |
| minimax-h3, minimax-h3-max | 15 s | 5 s |
| gemini-omni-flash-1.1 | 10 s | 3 s |
What does a 31-second video cost as two jobs?
Wan 3.0 at 720p bills $0.125 per second, so a 30-second job is $3.75 and a 2-second tail is $0.25: $4.00 in total, with the generation limit respected by both. At 1080p the same pair is $7.50 plus $0.50. Seedance 2.5 is priced per video token and costs more: about $17.34 for 30 seconds at 720p, before the tail.
Wan's smallest legal job is 2 seconds, so a 1-second tail does not exist. If you need exactly 31 seconds, generate 30 + 2 and cut one second off either clip with Video trim, which is $0.02 per job.
Will the two clips match visually?
Not automatically. Each job is generated on its own, so character, lighting and camera can drift at the seam. Two things help. Feed the same reference images to both jobs through input_references, or end the first clip and start the second from the same still using frame_images with last_frame and first_frame.
Use frame_images for the continuation: when both frame_images and input_references are sent, frame_images wins and the job runs as image-to-video.
How do you submit the pair?
Two requests, each with its own Idempotency-Key, so a retry does not create a second charge.
import os, requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def submit(key, seconds, prompt):
r = requests.post(
"https://api.sume.com/v1/videos",
headers={**H, "Idempotency-Key": key},
json={"model": "wan-3.0", "prompt": prompt,
"duration": seconds, "resolution": "720p"},
)
r.raise_for_status()
return r.json()["id"]
main = submit("clip-31-a", 30, "A drone flight over a coastal town at dawn")
tail = submit("clip-31-b", 2, "The same coastal town, drone slows and hovers")
print(main, tail)When is one 30-second job the better plan?
If the story fits in 30 seconds, take one job. Joining clips adds a cut and a second chance for the look to drift. The two-job route is for a length you cannot shorten, such as a fixed 31-second slot.
Sources
Related posts
More in Models
- 7-second AI video API: which Sume ids take 7 s and the price
Veo documents 4, 6 and 8 second clips, so 7 is not an option there. All twelve Sume ids accept 7 s; Omni Flash 1.1 at 360p is cheapest.
- 9-second AI video API: Omni takes 9 s, Veo does not
A 9 second clip fits Gemini Omni Flash 1.1 (3-10 s) and most Sume ids, but not Veo 3.1, which tops out at 8 s. Billed prices for 9 s inside.
- AI sound effects: Suno Sounds vs a music prompt
Suno Sounds V5.5 generates sound effects for 2 credits flat. Sume has no SFX endpoint; a Music Router prompt with no vocals gives ambient audio you must check.
- AI video releases and shutdowns in 2026: one dated timeline
Twelve dated events from Google, Luma, Runway, OpenAI and ElevenLabs, January to October 2026, each taken from a vendor page read on 2026-10-03.
Written by Sume