Check on-screen text in a Wan 3.0 clip: video frames at exact seconds

Wan 3.0 claims text rendering in 12 languages. Before you ship, pull PNG stills at the seconds the text appears with Sume video frames and read them yourself.

5 min readSume
All posts

To verify text that Wan 3.0 drew inside a clip, extract PNG stills at the seconds where the text is on screen and read them at full size. On Sume that is POST /v1/video-frames with at[] and format: "png"; if you omit max_edge, frames keep the source size, so small print stays readable. Alibaba's Wan 3.0 README (read 2026-10-05) lists text rendering in 12 languages as a feature; the list of languages is on that page. A feature claim is not a guarantee for your exact string, font and script, so you check each shipped line.

Pick the seconds

Text in a generated clip does not appear at a fixed time. If your prompt asks for a sign that comes into view around second six, sample second 5, 6, 7 and 8. Four stills cost little and cover the window. For a clip where text is on screen the whole time, three stills (start, middle, end) show whether the letters stay stable or morph.

  • Name each line of text you expect, and the language, in a checklist.
  • List the candidate seconds for each line.
  • Request all of them in one at[]; the docs show a list of times per request.
  • Open the PNGs, not the thumbnails, and compare letter by letter.
  • Record pass or fail per line, and keep the failed frames to explain the retry.

The request

Video frames takes one media.sume.com clip of your workspace plus one of at[] or fps. A submit always returns 202, so poll the GET. If one instant fails, that frame has url: null and the job still completes.

curl -X POST https://api.sume.com/v1/video-frames \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: text-check-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/wan-sign-clip.mp4",
    "at": [5, 6, 7, 8],
    "format": "png"
  }'

What a failed line means

If a letter is wrong, you have three choices, in order of cost. First, change the prompt: shorter strings with fewer glyphs survive better than long ones. Second, regenerate that one clip. Third, stop asking the model for the text and burn it yourself. The captions docs say cues with text, start and end burn authored text without speech-to-text, and that the style resolves by script: Korean text gets a Hangul-capable style, since the Latin default burns Korean as tofu. For other scripts, test a cue before you build a batch around it.

Text in a generated clip: who owns the pixels (read 2026-10-05)
ApproachText accuracyCost of a fix
Model draws the text (prompt only)Unverified, check every lineRegenerate the clip
Model draws, you check with video framesVerified for the seconds you sampledRegenerate or switch to cues
You burn text with caption cuesExactly your stringRe-run the caption job only
You overlay a still with Timeline composeExactly your imageRe-run compose only

A habit for batches

When you generate many clips with text in several languages, check one clip per script before running the rest. A script that renders badly will usually do so in every clip, so the early check saves the later cost. Keep a short sheet of pass and fail by language and fold it into the prompt guidance you reuse.

Mind the cost model. Video frames is billed by its own compute, not by the clip, and a failed or null frame does not fail the whole job. So the check is cheap compared with generating a second clip, and it is worth running before you spend on a batch of variants. If you need the same frame in a different size, rerun with a max_edge rather than resizing a screenshot yourself, since the extract works from the source pixels.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume