Wan 3.0 prompt guide for API users: prompt vs request fields

Write a Wan 3.0 prompt as a shot list with sound cues, and set length and size in request fields. What to put in the prompt and what to send to Sume's wan-3.0.

4 min readSume
All posts

Write the prompt as a short shot list with the sound you want, and put length, resolution and ratio in request fields, not in the prompt. Alibaba's Wan 3.0 guide describes multi-shot prompts of 4 to 6 seconds per shot with timestamps, and says dialogue and sound effects can be specified in the prompt.

Prompt advice below is from Alibaba's guide and fal's Wan 3 page, read 2026-09-29. What Sume accepts is from its Video generation docs.

What goes in the prompt?

Scene, camera and sound. fal's example prompts each state the framing, the lens or camera style, the light, small physical details and the ambient sound in one paragraph. For a longer clip, Alibaba's guide uses a timestamp per shot, with 4 to 6 seconds each, up to a 30-second narrative.

[0-5s] Wide shot: a ramen counter at night, steam rising, quiet room tone.
[5-10s] Close-up: chopsticks lift noodles, broth drips, a ladle clinks.
[10-15s] The cook slides the bowl forward and says "Careful, hot."

What goes in request fields instead?

Length, size and model. Alibaba's guide sets duration from 2 to 30 seconds, and Sume's docs list the same range for wan-3.0. Sume's request takes duration, resolution and aspect_ratio as separate fields, so a sentence such as "make it 12 seconds" in the prompt does not set the length.

Where each lever lives when you call wan-3.0 on Sume, read 2026-09-29.
LeverWhereNote
Shots and timingPromptTimestamps per shot
Dialogue and effectsPromptSpoken lines in quotes
Lengthduration2 to 30 whole seconds
Resolutionresolution480p, 720p or 1080p
Seed and exact sizeNot acceptedseed and size return 400

Can I pass Wan's own prompt options?

Not through Sume. Alibaba's API examples send prompt_extend, a prompt rewriter switch. Sume's docs say every v1 model lists no passthrough parameters and a non-empty provider.options returns 400 unsupported_parameter, so write the full prompt yourself.

What does a request look like?

One job with the shot list above.

curl -X POST "https://api.sume.com/v1/videos" \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: wan-prompt-001" \
  -d '{
    "model": "wan-3.0",
    "prompt": "[0-5s] Wide shot: a ramen counter at night, steam rising. [5-10s] Close-up: chopsticks lift noodles, a ladle clinks.",
    "duration": 10,
    "resolution": "720p",
    "aspect_ratio": "16:9"
  }'

How do I write the sound?

Say it in the same sentence as the action. fal's example prompts end with room tone or a sound: quiet room tone, the clink of ceramic on steel, street ambience. Alibaba's guide says the model produces voice, effects and music together, so a line of speech goes in quotes inside the shot it belongs to.

Keep each shot to one action. A 4 to 6 second shot with one move and one sound cue is easier to read than a shot that has three moves, and it maps to the timestamp layout Alibaba describes.

How do I keep a series consistent?

Reuse the same style sentence at the start of every prompt: camera, lens, light and palette. fal's examples open each prompt with the subject and then fix the framing and light in one clause, which makes a good template. Change only the action and the sound between clips.

For a character that must look the same across clips, use references rather than longer prose; see Wan 3.0 omni reference video.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume