Check on-screen text in a Wan 3.0 clip: video frames at exact seconds
Wan 3.0 claims text rendering in 12 languages. Before you ship, pull PNG stills at the seconds the text appears with Sume video frames and read them yourself.

To verify text that Wan 3.0 drew inside a clip, extract PNG stills at the seconds where the text is on screen and read them at full size. On Sume that is POST /v1/video-frames with at[] and format: "png"; if you omit max_edge, frames keep the source size, so small print stays readable. Alibaba's Wan 3.0 README (read 2026-10-05) lists text rendering in 12 languages as a feature; the list of languages is on that page. A feature claim is not a guarantee for your exact string, font and script, so you check each shipped line.
Pick the seconds
Text in a generated clip does not appear at a fixed time. If your prompt asks for a sign that comes into view around second six, sample second 5, 6, 7 and 8. Four stills cost little and cover the window. For a clip where text is on screen the whole time, three stills (start, middle, end) show whether the letters stay stable or morph.
- Name each line of text you expect, and the language, in a checklist.
- List the candidate seconds for each line.
- Request all of them in one
at[]; the docs show a list of times per request. - Open the PNGs, not the thumbnails, and compare letter by letter.
- Record pass or fail per line, and keep the failed frames to explain the retry.
The request
Video frames takes one media.sume.com clip of your workspace plus one of at[] or fps. A submit always returns 202, so poll the GET. If one instant fails, that frame has url: null and the job still completes.
curl -X POST https://api.sume.com/v1/video-frames \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: text-check-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/wan-sign-clip.mp4",
"at": [5, 6, 7, 8],
"format": "png"
}'What a failed line means
If a letter is wrong, you have three choices, in order of cost. First, change the prompt: shorter strings with fewer glyphs survive better than long ones. Second, regenerate that one clip. Third, stop asking the model for the text and burn it yourself. The captions docs say cues with text, start and end burn authored text without speech-to-text, and that the style resolves by script: Korean text gets a Hangul-capable style, since the Latin default burns Korean as tofu. For other scripts, test a cue before you build a batch around it.
| Approach | Text accuracy | Cost of a fix |
|---|---|---|
| Model draws the text (prompt only) | Unverified, check every line | Regenerate the clip |
| Model draws, you check with video frames | Verified for the seconds you sampled | Regenerate or switch to cues |
| You burn text with caption cues | Exactly your string | Re-run the caption job only |
| You overlay a still with Timeline compose | Exactly your image | Re-run compose only |
A habit for batches
When you generate many clips with text in several languages, check one clip per script before running the rest. A script that renders badly will usually do so in every clip, so the early check saves the later cost. Keep a short sheet of pass and fail by language and fold it into the prompt guidance you reuse.
Mind the cost model. Video frames is billed by its own compute, not by the clip, and a failed or null frame does not fail the whole job. So the check is cheap compared with generating a second clip, and it is worth running before you spend on a batch of variants. If you need the same frame in a different size, rerun with a max_edge rather than resizing a screenshot yourself, since the extract works from the source pixels.
Sources
Related posts
More in Media tools
- Coffee roaster brand video music: a 40-second lofi bed
Brief a warm 40-second lofi bed with vinyl texture for a coffee roaster brand video: one generation and a one-minute render, $0.225 on Sume.
- Coffee roaster holiday video: caption in your brand colour with design
Caption a roaster's holiday video in a brand colour with `design.colors.active`, then restyle twice without a second transcription: three looks for $0.60.
- Color-grade AI video on Sume: lut3d is off, tone filters stay
Sume video filter has a dim op, a crop op and an allowlisted ffmpeg filtergraph. lut3d is not on it. What a grade can and cannot do, per the docs.
- Comedy skit music: one bed and two stings with AI music
Score a 45-second comedy skit with one light bed and two short stings: three Music Router generations at $0.125 each plus a Timeline render, $0.475 in total.
Written by Sume