Wan 3.0 text in 12 languages: what Sume sends and what it bills
Alibaba says Wan 3.0 renders long text in 12 languages. On Sume you steer it through the prompt: no language field, 480p $0.0625 to 1080p $0.25 a second.

Alibaba's Wan 3.0 repository lists text rendering in 12 languages as a headline feature, and Sume exposes wan-3.0 in the Video Router. You get that text rendering by writing the exact on-screen words into the prompt; the documented Wan fields include no language setting, and billing is per output second by resolution, from $0.0625 to $0.25.
What the vendor claims
The repository says Wan 3.0 "renders long-form text across 12 languages" (read 2026-10-05). It does not, in the part we read, name the twelve languages or give an accuracy figure. Treat any list of languages you see elsewhere as unverified until you find it on an Alibaba page, and test the scripts you care about.
- Native 30 second generation and up to 20 reference assets, including documents and webpages.
- Instruction-based and reference-based editing, plus up to 12 sequential images with a cohesive style.
- Text rendering across 12 languages.
What you can set on Sume
The Sume request for wan-3.0 carries a prompt, resolution (480p, 720p, 1080p), duration (2 to 30 seconds), aspect_ratio (adaptive, 16:9, 4:3, 1:1, 3:4, 9:16), generate_audio, and reference URLs. Nothing in that list picks a text language, so the language is whatever you write in the prompt.
Prompt the text, do not hope for it
Quote the exact string you want on screen, in the script you want, and say where it appears. Keep each string short. Then check the first clip before you spend on the rest.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: wan-text-es-001" \
-d '{
"model": "wan-3.0",
"prompt": "Storefront at dusk. A neon sign reads exactly: Abierto hasta las 9",
"resolution": "720p",
"duration": 6,
"aspect_ratio": "16:9"
}'What it bills
Sume bills the provider list price times 1.25 for each output second. The Wan list is $0.05, $0.10 and $0.20 a second at 480p, 720p and 1080p (fal list, 2026-08-24 per the repo).
| Resolution | Provider list per s | Sume per s | 6 s clip | 30 s clip |
|---|---|---|---|---|
| 480p | $0.05 | $0.0625 | $0.375 | $1.875 |
| 720p | $0.10 | $0.125 | $0.75 | $3.75 |
| 1080p | $0.20 | $0.25 | $1.50 | $7.50 |
Pick the tier by type size
Small type needs pixels. A headline survives 480p; a paragraph of body copy in a non-Latin script may not. That is a judgment from how video resolution works, not a vendor figure, so run a short test before you commit to a tier. The related proof-clip approach is in the two-second proof clip post.
Set a text budget per scene
Alibaba markets the feature around long-form text, and the repository does not say how many characters count as long. Our advice is practical: set a text budget per scene. One short line per shot is a safe starting point, two lines is a test, and a full paragraph should come from your editor, not from the model.
Spelling is the second risk. Diacritics, right-to-left scripts and scripts with joined letters are where any image or video model slips first, and the repository gives no per-language figures. Generate a short clip per script you ship, read it frame by frame, and keep a note of which scripts passed. Reusing that note across campaigns is worth more than any general claim about 12 languages.
Test cheaply before you scale
A pass costs little at the low end. Wan 3.0 accepts clips from 2 seconds on Sume, so a 2 second proof clip at 480p is 2 x $0.0625 = $0.125. Twelve proof clips, one per language you plan to ship, cost $1.50. Compare that with finding a spelling error after a 30 second 1080p render at $7.50.
Because the prompt is the only language control, put the text in quotes, state the script name in plain words, and repeat the exact string in your own review notes. If a clip garbles the text, change one thing at a time: shorten the string, enlarge the described sign, or move to the next resolution tier.
When to use it
Choose Wan 3.0 for on-screen text when you need a single clip longer than 15 seconds, since most other catalog rows stop at 15 seconds. If you only need a few words and an edit pass, an overlay in your editor can be cheaper than regenerating a clip. Confirm the supported fields at any time in the Video Router docs and the Video generation docs.
Sources
Sources: the vendor feature list is from the Wan 3.0 repository (read 2026-10-05); every Sume number is from the Sume docs and repository pricing tables.
Sources
Related posts
More in Models
- Wan 3.0 inputs: text, image, video, audio on Alibaba vs Sume wan-3.0
Alibaba documents text, image, video and audio inputs for Wan 3.0, up to 30 s. Sume's wan-3.0 does text, image plus end frame, and references, 2 to 30 s.
- What is a full-duplex AI avatar, and when is a rendered clip enough?
A full-duplex avatar listens and talks at once, as in Tavus Griffin-Lite. If your viewers never talk back, a rendered Sume avatar clip does the job.
- Which AI video model for how many seconds: five length bands on Sume
Pick a Sume video model by clip length: 2 s, 3 to 4 s, 5 to 10 s, 11 to 15 s, and 16 to 30 s. Which ids accept each band, from the Video Router docs.
- Which aspect_ratio values do Sume video models list: 9:16 and 16:9
Sume docs list 9:16 and 16:9 for Auto and Gemini Omni Flash 1.1. Other models vary, so read capabilities from the catalog. What to send for TikTok or Reels.
Written by Sume