4-second AI video API: which Sume models take 4 s and the price
Eight Sume video ids accept a 4-second clip; MiniMax H3 does not. Price per 4-second job from $0.05 to $2.32 by model, with the 720p row for each.

A 4-second clip is the first length that almost every Sume video model accepts. Wan 3.0, Gemini Omni Flash 1.1, Kling 3, Grok Imagine Video 1.5 and the four Seedance ids all take it. MiniMax H3 and H3 Max do not, because they start at 5 seconds.
Priced at 720p, a 4-second job runs from $0.50 on Wan 3.0 or Omni up to $2.32 on Seedance 2.5.
What does a 4-second job cost on each model?
Sume bills the provider list price times 1.25 and rounds each job up to the cent. Per-second models are exact. Seedance is priced by video tokens, so its figures are the repo's estimate for a 16:9 clip with no reference video.
| Model id | 480p | 720p | 1080p | Note |
|---|---|---|---|---|
| grok-imagine-video-1.5 | $0.05 | $0.05 | n/a | Image-to-video only; flat rate |
| wan-3.0 | $0.25 | $0.50 | $1.00 | 2 to 30 s |
| gemini-omni-flash-1.1 | n/a | $0.50 | $0.75 | 360p is $0.15, 4K is $1.50 |
| kling-3 | n/a | $0.56 | $0.56 | $0.84 with audio on |
| seedance-2-mini | $0.36 | $0.76 | see catalog | Explicit pick only |
| seedance-2-fast | $0.57 | $1.21 | see catalog | Explicit pick only |
| seedance-2 | $0.71 | $1.52 | $3.41 | Explicit pick only |
| seedance-2.5 | $1.08 | $2.32 | $5.69 | 4 to 30 s |
Which 4-second row is the best value?
It depends on what the clip must do. Grok Imagine Video 1.5 is the cheapest, but it needs a starting image (image_url or a first frame) and takes no end frame or references. Wan 3.0 and Omni match at 720p and both make sound; on Omni the audio is always on and generate_audio: false is rejected.
Kling 3 takes no reference images, videos or audio. If the shot needs a character held across clips, Wan, Seedance and the MiniMax rows accept references and Kling does not.
Why is the 4-second Seedance price so much higher?
Seedance bills by video tokens, which grow with width, height, seconds and frame rate. The 2.5 model also uses a different token rate at 1080p. A 4-second 720p Seedance 2.5 clip is about $2.32; the same four seconds on Wan 3.0 is $0.50. What you buy for the difference is the 30-second ceiling and the reference limits, not the first four seconds.
How do you pin a 4-second request?
Send the model id and duration: 4 to POST /v1/videos. With sume/auto a 4-second request is valid because Auto accepts 3 to 10 seconds, but the family that serves it is not disclosed, so pin a model when you need the figures in the table.
import os, requests
r = requests.post(
"https://api.sume.com/v1/videos",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
json={
"model": "kling-3",
"prompt": "Rain on a neon street, slow push-in",
"duration": 4,
"resolution": "720p",
"aspect_ratio": "16:9",
},
)
print(r.status_code, r.json()["id"])What should you check before relying on the table?
Prices move with provider lists. Read pricing_skus and supported_durations from GET /v1/videos/models, or compare the live billed amount in the poll response's usage.cost after the first job.
Sources
Related posts
More in Pricing
- 40-second AI video API: a 30 s job plus a 10 s job on Sume
A 40-second AI video on Sume is a 30-second job plus a 10-second job. Wan 3.0 costs $2.51, $5.00 or $10.00 at 480p, 720p and 1080p. Split, price, join.
- 45-second AI video API: one 30 s job plus one 15 s job on Sume
A 45-second AI video on Sume is a 30-second Wan 3.0 job plus a 15-second one. Cost at 480p, 720p and 1080p, and which pairing keeps audio and look consistent.
- 5-second AI video API: every Sume model that takes it, priced
Ten Sume video ids take a 5-second clip, priced per job from $0.07 to $7.11. Per-model price at the lowest, 720p and highest tier, read 2026-10-04.
- 50-second AI video API: 30 + 20 seconds on Wan 3.0, cost by tier
A 50-second AI video on Sume takes two Wan 3.0 jobs, 30 s and 20 s: $3.13 at 480p, $6.25 at 720p, $12.50 at 1080p. Why Wan is the only model that can do it.
Written by Sume