Vertical 9:16 text-to-video API: a 12-second Wan 3.0 clip, priced

A 12-second 9:16 text-to-video request on Sume with Wan 3.0 costs $0.75 at 480p, $1.50 at 720p and $3.00 at 1080p. Full body and the other models that fit.

4 min readSume
All posts

A 12-second vertical text-to-video clip on Sume with Wan 3.0 bills $0.75 at 480p, $1.50 at 720p, and $3.00 at 1080p. Wan 3.0 accepts 2 to 30 seconds and lists 9:16, so 12 seconds is inside its range. Gemini Omni Flash 1.1 stops at 10 seconds, so it cannot make this clip in one job.

The request

This is the Video Router body from the docs, with the model changed to Wan 3.0 and a 720p tier. mode: "async" returns a job envelope you then poll.

curl -X POST https://api.sume.com/v1/video-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: vertical-12s-001" \
  -d '{
    "model": "wan-3.0",
    "prompt": "A vertical UGC-style product clip on a desk, natural light",
    "resolution": "720p",
    "duration": 12,
    "aspect_ratio": "9:16",
    "mode": "async"
  }'

Price by tier

Sume bills list times 1.25, rounded up to the cent per job. At 12 seconds the numbers divide evenly, so rounding adds nothing here.

12-second Wan 3.0 clip price (catalog rates, read 2026-10-05)
ResolutionRate per second12 s clip
480p$0.0625$0.75
720p$0.125$1.50
1080p$0.250$3.00

Other models that take 12 seconds in 9:16

Seedance 2.5 accepts 4 to 30 seconds and Seedance 2.0 up to 15, both with 9:16. The MiniMax H3 models accept 5 to 15 seconds. Each reports its own aspect ratios, so read supported_aspect_ratios first. Wan 3.0 lists auto, adaptive, 16:9, 4:3, 1:1, 3:4 and 9:16 in the catalog, and does not list 21:9.

If you want Sume to pick, model: "sume/auto" on /v1/videos takes 3 to 10 second clips at 16:9 or 9:16, so a 12-second request needs a pinned model.

After the submit

mode: "async" returns a job. Poll GET /v1/jobs/{id}/status or the polling URL, and fetch the result with GET /v1/jobs/{id}/result when result_ready is true or the status is completed. Completed jobs return artifacts under media.sume.com. The docs say not to treat queued as a failure: the job may be waiting for a concurrency slot.

If your workspace is on the Free plan, a 12-second job at 1080p still renders alone, because Free runs one job at a time. For a set of vertical clips, plan the waves from the plan's concurrency, not from the price.

Before you ship anything, read the live pages again: the catalog is public, the pricing page is public, and the docs describe the request fields. A blog post is a snapshot. The catalog, the plan grid and the error table are the things that change, so write your code to read them instead of copying numbers from a page, and re-check when a new model is added.

A good habit is a small log line per submit with the model, resolution, duration, estimated cost, job id and the Idempotency-Key you used. When a job misbehaves, those six fields answer most of the questions support will ask, and they let you compare your estimate with usage.cost and the usage ledger without re-running anything.

If you are new to the API, start with one clip, one model and the lowest resolution, read the full response once, and only then build a loop around it. Most surprises with video jobs come from fields that were defaulted, not from fields that were set.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume