Vidu Q4 audio references vs a scripted avatar: which keeps a voice?
Vidu Q4 Preview takes up to 3 audio references; Sume does not list Vidu. How a scripted avatar video keeps a presenter's face and words consistent.
Vidu Q4 Preview, announced on October 7, 2026, accepts up to three audio references to keep voices consistent across shots. Sume does not list a Vidu model, so you cannot call it through Sume. If your goal is one presenter who looks and sounds the same in every video, Sume's route is a reusable avatar plus a script.
What Vidu announced
Facts below come only from the vendor's press release (Vidu Q4 Preview on PR Newswire, read 2026-10-08). The release gives no lip-sync specifics, so this post makes no claim about how well audio references sync to mouth movement.
| Item | Vidu release, read 2026-10-08 |
|---|---|
| Announced | 2026-10-07, on the Vidu web product and API |
| Audio references | Up to 3, for voice consistency |
| Image references | Up to 15 |
| Output sizes | 540p, 720p, 1080p, 2K and 4K, 10-bit color |
| Launch price | $0.014 per second |
Does Sume list it?
No. The Sume model pages under Models list Avatar 1.0, Image, Video, Music and utilities; Vidu is not among them. Do not plan a Sume workflow that depends on Vidu, and do not infer a Sume price from the vendor price above.
How Sume keeps a presenter consistent
Consistency in Avatar 1.0 comes from identity, not from audio samples. You create an avatar once from a prompt, a structured profile or a photo, store it under a handle, and reuse that handle for every video. The words come from your script, so the content is as consistent as your copy.
Per the Generate avatar video page, one video resolves to one avatar, and multi-scene video_inputs are expected to share one scene. That is a constraint, not a gap: it keeps the face steady within a clip.
- Create once: POST /v1/avatar-1.0/generate.
- Reuse the handle: the leading @ is normalized and stored without it.
- Vary only the script, scene prompt or product_image between videos.
- Use quality plus by default, or max when quality matters more than turnaround.
When to pick which
Pick Vidu if you need a general video model with many references and are happy to use its own web product or API. Pick Sume Avatar 1.0 if the deliverable is a spokesperson speaking a script. Check the vendor's current page before relying on any figure above, because launch pricing is promotional by nature.
Reading a launch announcement carefully
A press release is the vendor's framing. Vidu's states launch pricing at $0.014 per second, and promotional pricing can end. It states up to three audio references, but not how audio is matched to a mouth in a talking shot. Before you build around it, run a short test of your own with the vendor's product and read its current documentation.
On the Sume side, nothing about Vidu changes what Avatar 1.0 does. If you need a face that speaks your exact script, use the talking-video route. If you need a recorded voice on a still, use the Fabric route. If you need cinematic motion with no speech, use the video router and its listed models. Keeping those jobs separate is the most reliable way to avoid mismatched lips.
If Sume later lists a Vidu model, it will appear in the Models documentation. Until then, assume it is not available through Sume.
Sources
Related posts
More in Sume Avatar 1.0
- What Sume Avatar 1.0 does not do: eight limits to check first
No streaming, no interruption, English-only speech in code, 720p, 4 to 60 seconds, one avatar per video. The limits of Sume Avatar 1.0 in one table.
- YouTube AI disclosure: which avatar pipeline steps are exempt?
YouTube exempts scripts, captions and upscaling from its AI label but not realistic synthetic people. Map each step of a Sume avatar pipeline to the rule.
- Does a scripted AI presenter video need YouTube's synthetic label?
YouTube asks for a label when content makes a real person appear to say something they did not. Here is how to read that for a made-up Sume Avatar presenter.
- Will a YouTube AI label hurt reach? What the page says
YouTube's help page says disclosing AI content does not limit reach or monetization eligibility; penalties target non-disclosure. Plan an avatar series on that.
Written by Sume