Vidu Q4 reference voice: audio references or TTS on Sume?

Vidu Q4 Preview adds reference voice. Sume does not list Vidu; some video rows take audio references, and TTS costs $0.0475 per 1,000 characters.

5 min readSume
All posts

The short answer

Vidu Q4 Preview lists reference voice as a launch feature, meaning a clip's voice can follow a sample. Sume does not list Vidu. The nearest options on Sume are two: send an audio reference to a video row that accepts one, or write the line as text and generate speech separately. The Sume docs say Seedance 2.x, Wan 3.0, MiniMax H3 and MiniMax H3 Max accept audio references. Gemini Omni Flash 1.1 does not; its native audio is always on and it rejects generate_audio: false.

Two routes compared

The docs list which rows accept audio, but they do not describe an audio reference as voice cloning, so test before you rely on it for a recurring character. The table sets the two routes side by side for a 6-second clip with one 300-character line.

Two ways to get a consistent voice on Sume, as of 2026-10-08 (Sume docs and catalog, read 2026-10-08)
RouteWhat you sendCost of the example
Audio reference on a video rowreference_audio_urls plus the promptWan 3.0 at 720p: 6 x $0.125 = $0.75
Audio reference on a video rowsame, on Seedance 2.0 at 720p6 x $0.378 = $2.268
Separate TTS300 characters of text300 / 1,000 x $0.0475 = $0.01425
Separate TTS, then timeline renderthe speech plus a silent clip$0.01425 + $0.10 minimum render ($0.10 per output minute, minimum 1 minute) plus the clip

Which to pick

Use an audio reference when the voice and the picture need to move together, for example a character speaking on camera. The model sees both in one pass, but you have little control over timing.

Use TTS plus a timeline when the line is fixed, the timing must be exact, or the same voice has to run across ten clips. The text is yours to edit, and a rerun of one line costs a cent or two. Omni has no audio reference, so an Omni-based pipeline needs this route anyway.

A third option is an avatar. If the speaker is a presenter rather than a character in a scene, Sume's avatar video route takes a script and a stored avatar, which keeps one voice and one face for every video. It costs $0.184 per second at the standard tier, so a 6-second clip is 6 x $0.184 = $1.104. The avatar comparison covers when that fits.

What to check before you commit

Run three clips with the same reference and the same line, and listen for drift across them. Vidu's release does not publish a limit for how many voice samples it takes or how long they may run, so read the vendor's own docs for that. On Sume, read each row's supported_input_references from GET /v1/videos/models in the video generation docs so you know if audio_url is accepted. The audio reference comparison lists the rows.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume