Four characters, two locations: a Wan 3.0 series bible in ten images

Wan 3.0 on Sume takes 10 reference images per job. Spend them as four characters at two views each plus two sets, and reuse the same set for every episode.

5 min readSume
All posts

With ten reference images per wan-3.0 job on Sume, a workable series bible is four characters at two views each (front and three-quarter) plus two location images. Reuse that exact set in every episode job. Ten 30-second episodes at 720p then cost 10 x $3.75 = $37.50, and the bible itself costs nothing in generation because the images are your own.

A ten-image Wan 3.0 reference set (read 2026-10-07)
SlotCountContent
Character front view4Neutral light, plain background
Character three-quarter view4Same outfit, same light
Location2The two sets used most

What the vendors claim and what you control

Alibaba says reference-to-video mode keeps characters, props, spaces and style consistent, including hairstyle, clothing and accessories. Morphic's page agrees and adds that labelling each reference in the prompt matters as much as supplying it. Both are vendor statements. On Sume you control the reference set and the prompt; consistency is something you test, not assume.

How to spend ten slots

  • Do not use two different outfits for one character in the same set; the model has no way to know which applies when.
  • Photograph or generate all character images with the same light so that differences come from the person.
  • Drop a character from the set in episodes where they do not appear. Fewer references leaves room for a third view of the lead.
  • Keep the set order the same across episodes and describe the characters in the same order in the prompt.

Episode costs

Thirty seconds at 720p is $3.75. At 480p it is $1.88 and at 1080p $7.50. A ten-episode season at those tiers is $18.80, $37.50 and $75.00. A single 480p test of the reference set costs $1.88 and is worth running before episode one.

Testing the set before episode one

Run the cheapest possible test first: one 480p job at 6 s, which is 6 x $0.0625 = $0.375, billed as $0.38. Ask for each character to turn toward the camera in turn. If one of the four looks different from the reference, replace that character's images before you spend $37.50 on a season. A small test also shows whether the two location images are being used or ignored.

Keep a written note of the exact prompt wording for each character, because the same wording in each episode helps. Sume's docs document no tag syntax for references, so descriptions in plain words are the only link between a name in the script and an image in the set.

Where it breaks

Consistency drifts when a prompt asks for a close-up of a character shown only from far away in the references, or for a back view that none of the images show. Sume's docs say references act as visual guidance rather than exact frames, so expect small changes in face and clothing detail. For a shot where the face must be exact, use a first frame instead, which makes the job image-to-video and means references are not used for that job.

A last frame from one episode can also open the next. Pull it with video-frames at a time just below the clip's duration, and send it as a first_frame for the next job. Because frame_images wins over references, that episode cannot also use the bible, so keep this for cold opens that continue a scene.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume