Wan 3.0 same-face test: eight 480p takes for $3.04 before 1080p

Alibaba says Wan 3.0 avoids same-face AI people. Test that claim on Sume with eight 6-second 480p takes at $0.38 each, then render only the best at 1080p.

5 min readSume
All posts

Alibaba's launch post says Wan 3.0's text-to-video mode is built for diverse, lifelike faces, with "no more same-face AI people". That is a vendor claim, and you can check it yourself on Sume for $3.04: eight 6-second Wan 3.0 jobs at 480p, one still from each, laid side by side. If the faces are distinct, spend $1.50 on a 1080p render of the best one.

The test needs no reference images, only eight prompts that describe different people. Treat the 480p result as a screening pass and look again at the final render, because the 1080p job is a separate generation.

Cost of the test

Wan 3.0 at 480p lists at $0.05 per second. Sume bills the list price times 1.25, which is $0.0625 per second. A 6 s take is $0.375, rounded up to $0.38. Eight takes are 8 x $0.38 = $3.04. The final 6 s at 1080p is 6 x $0.20 x 1.25 = $1.50. The whole exercise is $4.54, plus the Modal compute for the still extraction.

Cost of a Wan 3.0 same-face test on Sume (read 2026-10-07)
StepJobPrice on Sume
Eight screening takeswan-3.0, 480p, 6 s, no references$3.04
One finished takewan-3.0, 1080p, 6 s$1.50
Stills from each takevideo-frames, one frame at 3 sModal compute, billed per job

Running it

Submit each take to POST /v1/videos with model wan-3.0, resolution 480p, duration 6, and a prompt that fixes everything except the person: same street, same lens, same lighting, and a different age, build and hair each time. Add an Idempotency-Key per take so a retry does not queue a second job.

To pull stills, import each finished clip with POST /v1/media-imports, because video frames only reads media.sume.com artifacts. Then call POST /v1/video-frames with at set to [3]. The value must be at least 0 and below the clip's probed duration, so 3 is safe for a 6 s clip. The result is a durable image per take.

A prompt template that isolates the face

Write one fixed sentence for the scene and one variable sentence for the person. For example, the fixed part is a kitchen at 8 a.m., a 35 mm look, a single locked-off shot. The variable part names age range, build, hair and clothing. Keep the speech and action identical, such as turning toward the camera and smiling slightly, so the only difference between takes is the person.

Six seconds is long enough for a turn and a reaction, and short enough that the 480p tier keeps the bill under a dollar for any pair of takes. If you want sound, leave generate_audio at its default; if you do not, set it to false on both the screening and final jobs so the comparison stays like for like.

How to judge the grid

  • Compare the eight stills for repeated jaw, eye spacing and skin tone, not just hair.
  • Run the eight prompts twice. If the second pass keeps returning the same few faces, the prompts are not producing variety.
  • Check hands and background faces too: the vendor page itself lists audio texture and on-screen text accuracy as still improving, so do not assume other fidelity claims are settled.
  • Reference-to-video is a different test. Reference images are visual guidance on Sume, not exact frames, so judge that mode with its own grid.

Sources

Related posts

More in Models

All Models posts

Written by Sume