Wan 3.0 API in Node: vertical image-to-video from a first frame
Call Wan 3.0 on Sume from Node with fetch: pin a first frame, ask for 9:16 at 720p, poll the job and save the MP4. Includes limits and the 8-second price.

To use Wan 3.0 by API on Sume, POST to https://api.sume.com/v1/videos with model: "wan-3.0", put your still in frame_images as a first_frame, then poll the returned polling_url until the status is completed and download unsigned_urls[0]. The script below does exactly that with Node's built-in fetch.
An 8-second 9:16 clip at 720p costs $1.00 on Sume (provider list $0.10 per second times 1.25).
The script
Save it as wan.mjs and run SUME_API_KEY=... node wan.mjs on Node 18 or later. It uses the OpenRouter-shaped wire that /v1/videos follows, so submit and poll answer with bare objects, not a data envelope.
const H = { Authorization: `Bearer ${process.env.SUME_API_KEY}`, "Content-Type": "application/json" };
const body = {
model: "wan-3.0",
prompt: "Slow dolly in, steam rises from the cup, soft morning light",
duration: 8,
resolution: "720p",
aspect_ratio: "9:16",
frame_images: [{ type: "image_url", image_url: { url: "https://example.com/cup.png" }, frame_type: "first_frame" }],
};
async function main() {
const sub = await fetch("https://api.sume.com/v1/videos", { method: "POST", headers: H, body: JSON.stringify(body) });
if (sub.status !== 202) throw new Error(`${sub.status} ${await sub.text()}`);
const job = await sub.json();
for (;;) {
const s = await (await fetch(job.polling_url, { headers: H })).json();
if (s.status === "completed") {
const res = await fetch(s.unsigned_urls[0], { headers: H });
await (await import("node:fs/promises")).writeFile("wan.mp4", Buffer.from(await res.arrayBuffer()));
return console.log("saved wan.mp4, cost", s.usage?.cost);
}
if (s.status === "failed" || s.status === "cancelled") throw new Error(s.error ?? s.status);
await new Promise((r) => setTimeout(r, 10000));
}
}
main();What Wan 3.0 accepts on Sume
The catalog row gives Wan 3.0 a 2 to 30 second range, 480p, 720p and 1080p, and text, image, first and last frame, and reference inputs. Alibaba Cloud's Model Studio page lists the same three resolutions and a 30-second maximum for the hosted model, and describes text, image, video and audio inputs (read 2026-10-07).
| Field | Limit |
|---|---|
| duration | 2 to 30 whole seconds |
| resolution | 480p, 720p, 1080p (default 1080p when omitted) |
| aspect_ratio | Read supported_aspect_ratios from GET /v1/videos/models |
| input_references | Up to 10 images, 5 videos (15 seconds in total, 16 fps or more), 5 audio clips (15 seconds in total) |
| seed, size, provider.options | Rejected with 400 unsupported_parameter |
Things that trip first runs
The default resolution is 1080p, which is the most expensive tier, so set resolution on purpose. Document and web-page inputs (file_url, web_url) and enable_thinking are not exposed in v1, even though the vendor model supports documents.
If both frame_images and input_references are present, frame_images wins and the job runs as image-to-video. A URL that Sume cannot download fails the job with a message that names the media type, not the URL. Add an Idempotency-Key header if your code retries a submit, so a replay returns the original job instead of a second charge.
- Submit answers
202withid,polling_url,statusandmodel. - Statuses are
pending,in_progress,completed,failed,cancelled. - Prefer
callback_urlover polling for long jobs; it must be HTTPS.
Sources
Related posts
More in Developers
- Wan 3.0 API request cheat sheet: three modes, 2 to 30 seconds
Wan 3.0 on Sume: the request body for text, first/last frame and reference modes, the 480p/720p/1080p rates and the 2 to 30 second window, on one page.
- Wan 3.0 API rate limit: 300 RPM on Model Studio vs Sume jobs
Alibaba lists 300 requests per minute for Wan 3.0 on Model Studio. Here is what that means for a batch, and how Sume submits Wan 3.0 as async jobs.
- Hackathon app on the Sume Free plan: 1 seat, 6 accepted jobs
A weekend demo on Free can have one job processing and five queued. How to design the UI, the retries and the demo script around that, with the real limits.
- GPT Image 1 to GPT Image 2.5 on Sume: what changes in the output
Moving from GPT Image 1 to ChatGPT Image 2.5 on Sume changes the response (URL, not base64), default quality, size grid and failures.
Written by Sume