Vidu Q4 audio references vs a scripted avatar: which keeps a voice?

Vidu Q4 Preview takes up to 3 audio references; Sume does not list Vidu. How a scripted avatar video keeps a presenter's face and words consistent.

5 min readSume
All posts

Vidu Q4 Preview, announced on October 7, 2026, accepts up to three audio references to keep voices consistent across shots. Sume does not list a Vidu model, so you cannot call it through Sume. If your goal is one presenter who looks and sounds the same in every video, Sume's route is a reusable avatar plus a script.

What Vidu announced

Facts below come only from the vendor's press release (Vidu Q4 Preview on PR Newswire, read 2026-10-08). The release gives no lip-sync specifics, so this post makes no claim about how well audio references sync to mouth movement.

Vidu Q4 Preview facts, as of 2026-10-08
ItemVidu release, read 2026-10-08
Announced2026-10-07, on the Vidu web product and API
Audio referencesUp to 3, for voice consistency
Image referencesUp to 15
Output sizes540p, 720p, 1080p, 2K and 4K, 10-bit color
Launch price$0.014 per second

Does Sume list it?

No. The Sume model pages under Models list Avatar 1.0, Image, Video, Music and utilities; Vidu is not among them. Do not plan a Sume workflow that depends on Vidu, and do not infer a Sume price from the vendor price above.

How Sume keeps a presenter consistent

Consistency in Avatar 1.0 comes from identity, not from audio samples. You create an avatar once from a prompt, a structured profile or a photo, store it under a handle, and reuse that handle for every video. The words come from your script, so the content is as consistent as your copy.

Per the Generate avatar video page, one video resolves to one avatar, and multi-scene video_inputs are expected to share one scene. That is a constraint, not a gap: it keeps the face steady within a clip.

  • Create once: POST /v1/avatar-1.0/generate.
  • Reuse the handle: the leading @ is normalized and stored without it.
  • Vary only the script, scene prompt or product_image between videos.
  • Use quality plus by default, or max when quality matters more than turnaround.

When to pick which

Pick Vidu if you need a general video model with many references and are happy to use its own web product or API. Pick Sume Avatar 1.0 if the deliverable is a spokesperson speaking a script. Check the vendor's current page before relying on any figure above, because launch pricing is promotional by nature.

Reading a launch announcement carefully

A press release is the vendor's framing. Vidu's states launch pricing at $0.014 per second, and promotional pricing can end. It states up to three audio references, but not how audio is matched to a mouth in a talking shot. Before you build around it, run a short test of your own with the vendor's product and read its current documentation.

On the Sume side, nothing about Vidu changes what Avatar 1.0 does. If you need a face that speaks your exact script, use the talking-video route. If you need a recorded voice on a still, use the Fabric route. If you need cinematic motion with no speech, use the video router and its listed models. Keeping those jobs separate is the most reliable way to avoid mismatched lips.

If Sume later lists a Vidu model, it will appear in the Models documentation. Until then, assume it is not available through Sume.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume