Wan 3.0 dialogue prompts: says, Lip sync, Voiceover, No dialogue
How to write spoken lines for Wan 3.0 clips: the prompt grammar Alibaba documents, plus what the Sume /v1/videos request and price do with audio.

To make a Wan 3.0 clip speak, write the line inside the prompt as a character, a speaking verb, a colon and quoted text, for example she whispers: "Come closer.". Add Lip sync for mouth movement and write No dialogue when you want silence. Sume passes the prompt through unchanged and bills audio and silent clips the same.
The grammar Alibaba documents
The Wan3.0 prompt guide gives a small set of patterns. They are plain prompt text, not request fields, so the same text works from any client that can send a prompt.
- Single speaker: character, speaking verb, colon, quoted content.
- Several speakers with reference images:
Image 1 says: "...", thenImage 2 replies: "...". - Voice timbre from a reference clip: append
Voice timbre references Audio 1, or writeSpeak in the timbre of Audio 1: "...". - Natural mouth movement: add
Lip sync. - Narration with no on-screen speaker:
Voiceover: "...". - No speech at all: write
No dialogue, because the model may otherwise invent speech.
What Sume adds on top
On Sume the model id is wan-3.0 on POST /v1/videos. In the repo's pricing code generate_audio defaults to true and does not change the price, so a clip with spoken lines costs the same as a silent one. Price is the fal list per second times 1.25, and the reserve is held before the job runs. See the sound toggle post for the parameter itself.
Because audio is on by default, a prompt that says nothing about speech can still come back with generated sound. Put No dialogue in the prompt for clips you plan to score yourself.
Request with a spoken line
Keep the quoted line short enough to fit the clip. The prompt guide states no character limit for dialogue, so judge by length: a sentence that takes longer to say than the clip runs will be cut or rushed.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"wan-3.0","prompt":"A baker at a counter looks into the camera and says: \"Fresh loaves at six.\" Lip sync. Warm morning light.","duration":5,"resolution":"720p"}'When to use a voiceover line versus separate speech
A Voiceover: line keeps everything in one generation, which is simplest for a single clip. If a script spans many clips, a fixed narrator voice is easier to keep consistent from a separate speech step joined in Timeline, since each Wan generation interprets the timbre afresh unless you pass an audio reference. Pick by how many clips share the voice.
Check before you publish
Generation is not deterministic and seed is rejected on /v1/videos, so re-run if a line is garbled. Listen to the clip, or run video-inspect with transcribe: true to read back what was said; the docs list that at $0.01 per audio minute.
Worked examples
Here are three prompts built from the documented patterns. Each one states who speaks, what they say and whether mouth movement matters. Swap in your own subject; the structure is what the guide describes.
For a single speaker: A chef in a steamy kitchen turns to the camera and says: "Taste this." Lip sync. For two reference images: Image 1 says: "Did you hear that?" followed by Image 2 replies: "Hear what?". For narration over b-roll: Voiceover: "Ten minutes from the harbor." No on-screen speaker.
Keep each prompt to one idea per clip. If a scene needs five lines of conversation, split it across clips and join them afterwards rather than packing every line into one generation.
Common mistakes
Three mistakes account for most garbled results. Unquoted speech is treated as description, so the model may show someone talking without producing those words. Very long lines do not fit the clip length. And a prompt that asks for dialogue but also says nothing about audio references leaves the voice up to the model, which can differ from clip to clip.
- Always quote the spoken words.
- Match line length to clip length: a short sentence for a five second clip.
- Use an audio reference with the timbre phrasing when one voice must stay the same.
- Write
No dialoguefor clips meant to carry only music or ambience.
Budget the retries
Since seeds are not accepted, plan for more than one attempt on a line that matters. Because audio does not change the Wan price, retrying a spoken clip costs the same as retrying a silent one; the cost to watch is the resolution tier and the duration. Try the line at 480p first, then render the keeper at the resolution you need.
Sources
Related posts
More in Models
- 15 cuts in 30 seconds: Wan 3.0 2 s shots vs Seedance 2.5 4 s shots
A fast-cut 30 s ad on Sume: Wan 3.0 allows 2 s shots (15 cuts, $3.75 at 720p) while Seedance 2.5 and Kling 3 stop at 4 s (7 cuts). Prices side by side.
- Wan 3.0 reference videos on Sume: 5 clips, 15 s total, 16 fps floor
Wan 3.0 takes up to 5 reference videos, 10 images and 5 audios on Sume; other models stop at 3 videos. Catalog notes: 15 s combined, at least 16 fps.
- Wan 3.0 text in 12 languages: what Sume sends and what it bills
Alibaba says Wan 3.0 renders long text in 12 languages. On Sume you steer it through the prompt: no language field, 480p $0.0625 to 1080p $0.25 a second.
- Wan 3.0 inputs: text, image, video, audio on Alibaba vs Sume wan-3.0
Alibaba documents text, image, video and audio inputs for Wan 3.0, up to 30 s. Sume's wan-3.0 does text, image plus end frame, and references, 2 to 30 s.
Written by Sume