Wan 3.0 API request cheat sheet: three modes, 2 to 30 seconds

Wan 3.0 on Sume: the request body for text, first/last frame and reference modes, the 480p/720p/1080p rates and the 2 to 30 second window, on one page.

5 min readSume
All posts

Wan 3.0 on Sume is the catalog id wan-3.0. You call it on POST /v1/videos like any other video model: the mode is inferred from the fields you send, the length is duration in whole seconds from 2 to 30, and the resolution is 480p, 720p or 1080p. It supports text-to-video, image-to-video with an optional end frame, and reference-to-video, and it lists audio. This page keeps the three bodies and the rates together so you can copy one and go.

A good default for a first test is 5 seconds at 480p, which costs about $0.31 before rounding. Once the prompt works, raise the resolution and the length. Changing one thing at a time makes it clear what the change did, and the cheap tier is enough to judge the motion and the composition.

Read the live catalog with GET /v1/videos/models before you pin anything. The catalog row is the truth for a model and the numbers here come from the Sume video router documentation.

Which fields pick the mode

You never send a mode. Sume infers it from the body, and the rule is the same for every model on the route.

How the mode is chosen for wan-3.0 (Sume docs, read 2026-10-07)
You sendModeNotes
prompt onlyText-to-videoNeeds prompt, duration, resolution
frame_images with first_frameImage-to-videoPins the opening frame
frame_images with first_frame and last_frameImage-to-video with endwan-3.0 lists i2v plus end frame
input_referencesReference-to-videoConditions the whole clip; not a pinned opening frame
frame_images and input_referencesImage-to-videoframe_images wins

The text-to-video body

The smallest useful request has four fields. Set resolution explicitly, because the default differs between models, and set duration as an integer.

curl -X POST "https://api.sume.com/v1/videos" \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "wan-3.0",
    "prompt": "A tram crosses a rainy intersection at dusk, handheld camera",
    "duration": 8,
    "resolution": "720p",
    "aspect_ratio": "16:9"
  }'

First and last frame, and references

For image-to-video, add frame_images with the frame types you want. Send only a first_frame to start on your picture, or add a last_frame to land on a second picture. For reference-to-video, use input_references and remember the difference between the two: a reference conditions the clip but does not promise that the first frame matches it. If you need the clip to start on your photo, use frame_images, as explained in why a clip does not start on the photo.

Send an Idempotency-Key header on every submit, so that a retry after a dropped connection does not create a second job.

The submit call returns 202 with an id, a polling_url and a status. Poll that address, or pass an HTTPS callback_url in the body and receive a signed webhook when the job finishes. When it is complete, fetch the file from GET /v1/videos/{job_id}/content, which redirects to the artifact. A webhook is usually easier than a tight polling loop.

Rates and the 30 second ceiling

The list rates are $0.05, $0.10 and $0.20 per second at 480p, 720p and 1080p. Sume bills at the list rate times 1.25, rounded up to cents on the amount. A reservation is made when you submit, the amount is captured when the clip completes and the reservation is released if the job fails.

wan-3.0 price per clip, list x 1.25 before rounding up to cents (Sume docs, read 2026-10-07)
Length480p ($0.0625/s)720p ($0.125/s)1080p ($0.25/s)
2 s$0.125$0.25$0.50
5 s$0.3125$0.625$1.25
10 s$0.625$1.25$2.50
30 s$1.875$3.75$7.50

Choosing within the window

The 30 second ceiling is the largest in the catalog together with Seedance 2.5, and the 2 second floor is the lowest. If your shot is under 3 seconds, Wan 3.0 is one of the few that accepts it, since Omni starts at 3 and the rest at 4 or 5. For a map of every window, see duration windows per model, and for the capability fields, see reading capabilities on the models endpoint.

The practical rule is to pick the length from the shot, not from the model. Decide how long the action needs, then check that the number is in the 2 to 30 window, and send it as an integer. If the clip you want is longer than 30 seconds, plan two requests and cut between them at a natural break.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume