Wan 3.0 proof clip: test on-screen text before a 30-second render

A 2 second 480p Wan 3.0 clip costs $0.125 on Sume. Use it to test on-screen text per language before a 30 second render that costs $1.875 to $7.50.

5 min readSume
All posts

Before a long Wan 3.0 render, run a 2 second proof clip at 480p: it costs $0.125 on Sume, the 2 second minimum for wan-3.0, and shows whether your on-screen text renders. A 30 second render of the same prompt costs $1.875 at 480p, $3.75 at 720p and $7.50 at 1080p.

Why a proof clip

Alibaba's Wan 3.0 repository lists text rendering across 12 languages (read 2026-10-05) without per-language results in the part we read. Your script, your fonts in the prompt and your sign sizes decide what works. A proof clip is the cheapest way to find out.

What it costs

A proof clip uses the lowest tier and the shortest length, and holds the text constant.

Proof clips versus full renders on Sume's wan-3.0 (Sume pricing tables, read 2026-10-05)
ItemLength480p720p1080p
One proof clip2 s$0.125$0.25$0.50
Six-language proof set6 x 2 s$0.75$1.50$3.00
One full render30 s$1.875$3.75$7.50
Six full renders6 x 30 s$11.25$22.50$45.00

What it saves

A six-language proof set at 480p costs $0.75, which is 6.7 percent of the $11.25 for six full 480p renders. If two languages fail in the proof, you have saved two full renders you would otherwise throw away.

The proof prompt

Write the same short scene and put one quoted string in it. Change only the string and the language name. Read each result frame by frame, and write down the scripts that passed.

import asyncio, json, os, urllib.request

LINES = {"fr": "Ouvert tous les jours", "es": "Abierto todos los dias", "de": "Taglich geoffnet"}

def submit(lang, text):
    body = {"model": "wan-3.0", "resolution": "480p", "duration": 2,
            "aspect_ratio": "16:9", "generate_audio": False,
            "prompt": f"Shop window with a sign that reads exactly: {text}"}
    req = urllib.request.Request("https://api.sume.com/v1/videos",
        data=json.dumps(body).encode(), method="POST",
        headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
                 "Content-Type": "application/json",
                 "Idempotency-Key": f"proof-{lang}-v1"})
    return json.load(urllib.request.urlopen(req))["id"]

async def main():
    print(await asyncio.gather(*(asyncio.to_thread(submit, k, v) for k, v in LINES.items())))

asyncio.run(main())

What a proof does not prove

A proof clip tells you about the sign, not the whole film. Longer clips can drift from the start of the clip to the end, so a string that is correct at second 1 may change at second 25. The proof reduces the risk but does not remove it, and the final render still needs a full watch.

Overlay when exactness matters

When the text must be exact, consider overlay instead of rendering. Sume documents a standalone caption job, POST /v1/video-captions, which burns authored overlay cues (text with start and end times) onto a finished video URL, at $0.20 per accepted job for videos up to 60 seconds under the current estimate. Text that you place yourself is correct by construction, and the model only has to draw the picture.

What to look for in the frames

Check each frame for dropped or swapped letters, accents that move, and words that melt in the last frames of the clip. In a 2 second clip there are only about 48 frames at 24 per second, so scrubbing through them is quick. Note the scripts that fail and move those lines to overlay captions.

Keep the winning prompt string. When you scale to the 30 second render, change only the duration and resolution so that the result is comparable with the proof.

Proof by language group

Group your languages by script: Latin with accents, Cyrillic, Arabic, and the East Asian scripts. Run one proof per group at first, since a failure inside a group usually repeats across it. If the Latin proof passes and the Arabic proof fails, you have learned that at a cost of $0.25 at 480p, without testing all twelve.

Test the longest string you plan to use, not the shortest. Short words often render correctly while a long line of the same script drops letters, and the long line is the one on your final frame.

When to skip it

Skip the proof for a text-free film, for a template whose strings you have already verified, and when you will overlay all text with a caption job. A proof is a quality gate for new text, not a ritual for every render.

The routine, priced

Cost of the whole routine for six languages at 720p: six proofs at 480p ($0.75), then six 720p renders ($22.50), or $23.25. Without a proof, the same six renders cost $22.50, so a proof pays off only if it prevents a failed render. Run it for new scripts and layouts, skip it for a template you have already verified. Sources: the Wan 3.0 repository (read 2026-10-05), the Video Router docs and the Video generation docs.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume