Wan 3.0 API request cheat sheet: three modes, 2 to 30 seconds
Wan 3.0 on Sume: the request body for text, first/last frame and reference modes, the 480p/720p/1080p rates and the 2 to 30 second window, on one page.

Wan 3.0 on Sume is the catalog id wan-3.0. You call it on POST /v1/videos like any other video model: the mode is inferred from the fields you send, the length is duration in whole seconds from 2 to 30, and the resolution is 480p, 720p or 1080p. It supports text-to-video, image-to-video with an optional end frame, and reference-to-video, and it lists audio. This page keeps the three bodies and the rates together so you can copy one and go.
A good default for a first test is 5 seconds at 480p, which costs about $0.31 before rounding. Once the prompt works, raise the resolution and the length. Changing one thing at a time makes it clear what the change did, and the cheap tier is enough to judge the motion and the composition.
Read the live catalog with GET /v1/videos/models before you pin anything. The catalog row is the truth for a model and the numbers here come from the Sume video router documentation.
Which fields pick the mode
You never send a mode. Sume infers it from the body, and the rule is the same for every model on the route.
| You send | Mode | Notes |
|---|---|---|
| prompt only | Text-to-video | Needs prompt, duration, resolution |
| frame_images with first_frame | Image-to-video | Pins the opening frame |
| frame_images with first_frame and last_frame | Image-to-video with end | wan-3.0 lists i2v plus end frame |
| input_references | Reference-to-video | Conditions the whole clip; not a pinned opening frame |
| frame_images and input_references | Image-to-video | frame_images wins |
The text-to-video body
The smallest useful request has four fields. Set resolution explicitly, because the default differs between models, and set duration as an integer.
curl -X POST "https://api.sume.com/v1/videos" \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "wan-3.0",
"prompt": "A tram crosses a rainy intersection at dusk, handheld camera",
"duration": 8,
"resolution": "720p",
"aspect_ratio": "16:9"
}'First and last frame, and references
For image-to-video, add frame_images with the frame types you want. Send only a first_frame to start on your picture, or add a last_frame to land on a second picture. For reference-to-video, use input_references and remember the difference between the two: a reference conditions the clip but does not promise that the first frame matches it. If you need the clip to start on your photo, use frame_images, as explained in why a clip does not start on the photo.
Send an Idempotency-Key header on every submit, so that a retry after a dropped connection does not create a second job.
The submit call returns 202 with an id, a polling_url and a status. Poll that address, or pass an HTTPS callback_url in the body and receive a signed webhook when the job finishes. When it is complete, fetch the file from GET /v1/videos/{job_id}/content, which redirects to the artifact. A webhook is usually easier than a tight polling loop.
Rates and the 30 second ceiling
The list rates are $0.05, $0.10 and $0.20 per second at 480p, 720p and 1080p. Sume bills at the list rate times 1.25, rounded up to cents on the amount. A reservation is made when you submit, the amount is captured when the clip completes and the reservation is released if the job fails.
| Length | 480p ($0.0625/s) | 720p ($0.125/s) | 1080p ($0.25/s) |
|---|---|---|---|
| 2 s | $0.125 | $0.25 | $0.50 |
| 5 s | $0.3125 | $0.625 | $1.25 |
| 10 s | $0.625 | $1.25 | $2.50 |
| 30 s | $1.875 | $3.75 | $7.50 |
Choosing within the window
The 30 second ceiling is the largest in the catalog together with Seedance 2.5, and the 2 second floor is the lowest. If your shot is under 3 seconds, Wan 3.0 is one of the few that accepts it, since Omni starts at 3 and the rest at 4 or 5. For a map of every window, see duration windows per model, and for the capability fields, see reading capabilities on the models endpoint.
The practical rule is to pick the length from the shot, not from the model. Decide how long the action needs, then check that the number is in the 2 to 30 window, and send it as an integer. If the clip you want is longer than 30 seconds, plan two requests and cut between them at a natural break.
Sources
Related posts
More in Developers
- GPT Image 1 to GPT Image 2.5 on Sume: what changes in the output
Moving from GPT Image 1 to ChatGPT Image 2.5 on Sume changes the response (URL, not base64), default quality, size grid and failures.
- What is a partial transcript in streaming speech to text?
A partial is a provisional transcript a streaming model revises as audio arrives. Why subtitles for a finished clip only need final text and word times.
- What to show a viewer while an avatar video job is queued
Avatar jobs on Sume are async: queued, processing, then completed, failed or canceled. A status-to-UI map for waiting screens, with polling rules.
- When is async TTS the right choice? Sync wait, poll or webhook
Async TTS is right for voiceovers, batches and anything a person is not watching a spinner for. Sume's sync wait stops at 30 seconds; Flash claims 45 ms.
Written by Sume