12-second AI video API: nine Sume models that accept 12 seconds

Every prompt-driven row except Gemini Omni Flash 1.1 accepts 12 seconds on Sume. Per-second cost for Wan, MiniMax H3 and H3 Max, plus Kling's audio on and off.

5 min readSume
All posts

A 12-second request is accepted by 9 of the ten text-or-image video rows in Sume's catalog, and refused by the rest with a 400 at validation. Sume advertises each row's supported_durations at GET /v1/videos/models, and the Video Router docs list the same windows. Twelve seconds is just past Gemini Omni Flash 1.1's 10-second ceiling, which makes it a clean dividing line.

Which Sume models accept 12 seconds?

Accepted: Seedance 2.5 (seedance-2.5), Seedance 2.0 (seedance-2), Seedance 2.0 Fast (seedance-2-fast), Seedance 2.0 Mini (seedance-2-mini), Kling Video v3 Pro (kling-3), Wan 3.0 (wan-3.0), Grok Imagine Video 1.5 (grok-imagine-video-1.5), MiniMax H3 (minimax-h3), MiniMax H3 Max (minimax-h3-max).

Refused: gemini-omni-flash-1.1. For H3 Max Recast and Higgsfield Genjutsu the length is not a choice at all, since the output keeps the length of the source video.

  • Per-row windows, in whole seconds: seedance-2.5 (4 to 30 s), seedance-2 (4 to 15 s), seedance-2-fast (4 to 15 s), seedance-2-mini (4 to 15 s), kling-3 (4 to 15 s), wan-3.0 (2 to 30 s), grok-imagine-video-1.5 (4 to 15 s), minimax-h3 (5 to 15 s), minimax-h3-max (5 to 15 s), gemini-omni-flash-1.1 (3 to 10 s).

What does a 12-second clip cost?

Billable cost is the per-second rate times the seconds, rounded up to the cent, where the rate is the provider list times 1.25. The table prices only the rows with a published per-second rate; the Seedance rows bill per video token, so read their cost from the finished job rather than from this table.

The lowest and highest resolution columns bracket the range: resolution is the biggest lever on any per-second row, so a draft at the lowest tier and a final at the highest is the cheapest way to iterate.

12-second clip billable cost, Sume rate = provider list x 1.25, Sume docs (read 2026-10-03)
RowLowest resolutionHighest resolution
wan-3.0480p $0.751080p $3.00
minimax-h3480p $0.75768p $0.90
minimax-h3-max480p $0.751080p $2.40
kling-3720p, audio off $1.68audio on $2.52

How is a 12-second job billed and settled?

Sume reserves the workspace balance on submit at the provider list times 1.25 and reports the final billable amount as usage.cost on the poll response at GET /v1/videos/{id}. The reserve is an estimate from the request you sent, so a longer duration reserves more up front. If you send an Idempotency-Key, a retry of the same body returns the original job instead of creating and billing a second one.

Poll until the job reaches a terminal status, then download with GET /v1/videos/{id}/content?index=0. A content request on a failed job is a 409 job_failed, so check the status first. The Jobs and results page covers the lifecycle.

How do I request 12 seconds?

Send duration: 12 as a whole number on POST /v1/videos. Do not send a fractional value; the windows are in whole seconds. If you would rather not pick a row by hand, GET /v1/videos/models returns supported_durations, and filtering it for 12 is a few lines, as shown in the Python filter post.

Remember that sume/auto only validates 3 to 10 seconds, so a request outside that window needs a pinned model.

curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"seedance-2.5","prompt":"A slow dolly past a rain-streaked tram window","duration":12}'

Sources

Related posts

More in Models

All Models posts

Written by Sume