50-second AI video API: 30 + 20 seconds on Wan 3.0, cost by tier

A 50-second AI video on Sume takes two Wan 3.0 jobs, 30 s and 20 s: $3.13 at 480p, $6.25 at 720p, $12.50 at 1080p. Why Wan is the only model that can do it.

5 min readSume
All posts

A 50-second AI video on Sume is two jobs: 30 seconds and 20 seconds. Wan 3.0 is the model to use, because it accepts any whole number from 2 to 30 seconds, so a 20-second second half is legal. The pair costs $3.13 at 480p, $6.25 at 720p and $12.50 at 1080p.

Models capped at 15 seconds would need four jobs for 50 seconds.

Which models can make the 20-second half?

Only the rows with a maximum above 15 seconds can produce a clip of 16 to 30 seconds in one job.

Sume video models that accept a 20-second job, read 2026-10-04
Model idAccepts 20 sLongest
wan-3.0yes30 s
seedance-2.5yes30 s
seedance-2, seedance-2-fast, seedance-2-minino15 s
kling-3no15 s
grok-imagine-video-1.5no15 s
minimax-h3, minimax-h3-maxno15 s
gemini-omni-flash-1.1no10 s

What does 30 + 20 cost on Wan 3.0?

Each job is billed at the list rate times 1.25 and rounded up to the cent. A 30-second job at 720p is $3.75 and a 20-second one is $2.50.

  • 480p: $1.88 + $1.25 = $3.13
  • 720p: $3.75 + $2.50 = $6.25
  • 1080p: $7.50 + $5.00 = $12.50

How does it compare with four 15-second Kling jobs?

Kling 3 stops at 15 seconds, so 50 seconds means four jobs (for example 13, 13, 12 and 12 seconds), and Kling takes no reference images to hold a character across them. At $0.14 per second billed with audio off, 50 seconds of Kling is $7.00, close to Wan at 720p, but with three seams instead of one.

Fewer joins usually look better. Pick the model with the longest single clip that suits the look.

How do you request the pair?

Both halves are wan-3.0 jobs. The second takes a first frame cut from the end of the first.

import os, requests

H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

def job(key, secs, prompt, extra=None):
    body = {"model": "wan-3.0", "prompt": prompt,
            "duration": secs, "resolution": "720p", **(extra or {})}
    r = requests.post("https://api.sume.com/v1/videos",
                      headers={**H, "Idempotency-Key": key}, json=body)
    r.raise_for_status()
    return r.json()["id"]

print(job("film-50-a", 30, "A cyclist crosses a city at dawn"))
print(job("film-50-b", 20, "The cyclist reaches a quiet harbour"))

What should you check before you scale it?

Read usage.cost on the first completed 30-second job and compare with the table. If you plan many 50-second pieces, draft each half at 480p first.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume