Ad text in AI video: trust Wan 3.0 text, or fix the image first?
Alibaba says Wan 3.0 renders text in 12 languages. For a price or legal line, edit the text in a still with Ideogram 4.5, then animate it: $0.70 per language.

Trust model-rendered text for a mood line, and fix the text in a still image for a price or a legal line. Alibaba says Wan 3.0 renders text in 12 languages, but nothing on the vendor page promises exact characters on every frame, and a wrong digit in a sale price is a real cost.
The safer route is to settle the text in an image, where you can read it and fix it, and then animate that image. On Sume that is an Ideogram 4.5 edit at $0.075 for a medium-quality image, then a short image-to-video clip.
What each vendor claims
The Wan 3.0 repository (read 2026-10-05) lists text rendering in 12 languages. Ideogram's launch post (read 2026-10-05) calls 4.5 the most precise edit model, says it eliminates artifact buildup so that multi-turn editing is possible, and says it is live in Ideogram, the API and launch partners. Both are vendor claims about capability. Neither is a measured error rate, and I have none to offer.
On Sume, Ideogram 4.5 is ideogram/ideogram-v4.5 in the image catalog. Without input_references it generates from text. With references it edits the first image and uses up to four more as references, five in total. Quality is low, medium (default) or high, resolution is 1K or 2K, and an edit without aspect_ratio keeps the shape of the source image.
Which route for which text
The decision depends on what the text is.
| Ad element | Risk if the digits are wrong | Route |
|---|---|---|
| Brand tagline, mood word | Low | Let the video model render it, then review the frames |
| Sale price, discount percent | High | Edit it in a still, then animate the still |
| Legal or disclosure line | High | Burn it as a caption or end card you control |
| Translated headline | Medium | Edit the headline in the source image per language |
The cost of fixing text first
The two-step route is also cheap. Ideogram 4.5 on Sume is $0.0375, $0.075 or $0.275 per image by quality (list $0.03, $0.06, $0.22, times 1.25). A medium edit plus a 5-second Omni Flash 1.1 clip at 720p ($0.125 per second, $0.625) is $0.70 per language. Five languages cost $3.50.
Whichever route you choose, check the finished clip frame by frame for the price and the legal line before it goes live. Pull stills from the final MP4 and read them at full size, since a wrong character is easy to miss at playback speed.
For a deeper walk through the edit-then-animate route with code, see the translation pipeline post linked below.
Why fix text in a still
Start with the failure you can least afford: a price that is wrong in one language. If a model renders text inside a moving scene, the characters can change between frames, and you only discover it if someone reads every frame. A still image has one frame to read. That is the case for settling the text there first.
The second reason is control. When you edit a still, you give an instruction like 'change the headline to the Spanish line below and leave everything else alone', and you can compare the result against the source. Ideogram's own description of 4.5 is built around that loop: precise edits, with multi-turn editing possible. On Sume, an edit without aspect_ratio keeps the shape of the source, which is what you want when the still must match a 9:16 placement.
The third reason is cost shape. A bad still costs $0.075 to redo. A bad video costs more, for example $0.625 for a 5-second Omni clip at 720p, and re-rendering a 30-second Wan clip at 1080p costs $7.50.
Where model-rendered text is fine
Model-rendered text is not wrong in principle. For a wordless mood spot with one short brand word, it removes a step. For a product demo where the label on the box must read correctly, the label should come from the real product photo, not from generation. In both cases, keep the on-screen text short, in one language per clip, and review it at full size.
If you localize into many languages, make the text edit the first step per language and reuse the same animation prompt. That keeps the motion constant, so the only difference between the Spanish and German clips is the text, which makes the review quick.
Sources
Related posts
More in Comparisons
- How long can an AI video be in 2026? Max duration by model
Vendor pages list 20 to 40 seconds per generation in 2026. Sume's catalog ids accept 15 s for most, 30 s for three, and 10 s for Omni Flash 1.1.
- Alibaba Live Avatar needs five H800 GPUs: or call a hosted clip API
Alibaba-Quark's open Live Avatar streams audio-driven video at 45 FPS on five H800 GPUs. Compare that hardware with a hosted still-plus-audio clip.
- Anam plans from $12 to $999: cost per live minute vs a rendered one
Anam's plan ladder works out to $0.12 to $0.24 per included live minute. A 60-second Sume avatar clip costs $11.04, so it wins above about 74 viewers.
- Atmee avatar billed by the minute vs Sume reserve and refund
Atmee's LiveKit plugin bills by the minute, capped at 3,600 seconds by default. Sume reserves a clip price upfront and refunds failures. Compared.
Written by Sume