AI video ad text in 12 languages: Wan 3.0 or burned-in captions?
Alibaba says Wan 3.0 renders text in 12 languages. Sume lists wan-3.0 at 2 to 30 s. When to let the model draw words and when to add captions yourself.

If your ad needs words on screen in more than one language, you have two choices: let the video model draw the text, or add the text afterwards as captions. Alibaba's Wan 3.0 repository says the model renders text in 12 languages, and Sume's catalog lists wan-3.0. For anything where a misspelled word would cost you, add the text yourself.
What the vendor says and what Sume lists
The Wan 3.0 repository (read 2026-10-06) describes native 30 second generation, up to 20 reference assets, automatic scene splitting and text rendering in 12 languages. Sume's catalog code lists wan-3.0 at 2 to 30 seconds, with 480p, 720p and 1080p, and up to 10 image references, 5 video references and 5 audio references. Those reference limits differ from the vendor's, so follow Sume's numbers when calling Sume.
| Item | Alibaba page | Sume catalog |
|---|---|---|
| Length | Native 30 s | 2 to 30 s |
| Resolutions | Not the point here | 480p, 720p, 1080p |
| Reference assets | Up to 20 | Images up to 10, videos up to 5, audio up to 5 |
| Text rendering | 12 languages | Not a Sume claim |
Which approach for which text
Model-drawn text suits short, decorative words that are part of a scene, such as a sign or a label. Captions suit prices, dates, discount codes, legal lines and anything a customer will act on. A caption layer is exact because you typed it, and you can change a language without regenerating the video.
The 12 language claim is the vendor's. Sume does not verify it, and the docs do not promise accuracy per language. Test the exact words you need in each language, and look at every frame where text appears.
A workable setup
Generate the clip without relying on text for the message, then place the words on a Timeline 1.0 edit or in your own editor. Keep one master clip and one text file per language. If you do let the model draw a word, keep it to one or two and check it on a phone-sized screen before you ship.
Cost note
On pricing, read the pricing_skus field from GET /v1/videos/models rather than a number from a blog post. The cost of a 30 second clip depends on resolution, and a draft at 480p is the cheap way to test whether the model draws your words correctly before you spend on 1080p.
Sources
Related posts
More in Models
- AI video API news, October 6 2026: what changed for callers
Kling 4.0 heads to full launch, Seedance 2.5 API pages disagree, Omni 1.1 Flash is Google's default. What to change in a video integration today.
- Bumper sticker art: a 3:1 strip at 1536x512 on GPT Image 2.5
Make a wide bumper sticker on Sume with GPT Image 2.5 at a custom 1536x512 size, why 3:1 is its widest shape, and what to do for longer stickers.
- ChatGPT Images 2.5 Sketch: the API way is a drawing as a reference
ChatGPT Images 2.5 added a Sketch feature. On Sume there is no sketch canvas, but a drawing sent as an input reference to GPT Image 2.5 does the same job.
- Omni Flash or Veo 3.1 for the Gemini API? What Sume runs
Google's video docs now call Omni Flash the default over Veo 3.1. Here are Google's listed prices for both, and which one Sume's video router carries.
Written by Sume