Vidu Q4 reference voice: audio references or TTS on Sume?
Vidu Q4 Preview adds reference voice. Sume does not list Vidu; some video rows take audio references, and TTS costs $0.0475 per 1,000 characters.

The short answer
Vidu Q4 Preview lists reference voice as a launch feature, meaning a clip's voice can follow a sample. Sume does not list Vidu. The nearest options on Sume are two: send an audio reference to a video row that accepts one, or write the line as text and generate speech separately. The Sume docs say Seedance 2.x, Wan 3.0, MiniMax H3 and MiniMax H3 Max accept audio references. Gemini Omni Flash 1.1 does not; its native audio is always on and it rejects generate_audio: false.
Two routes compared
The docs list which rows accept audio, but they do not describe an audio reference as voice cloning, so test before you rely on it for a recurring character. The table sets the two routes side by side for a 6-second clip with one 300-character line.
| Route | What you send | Cost of the example |
|---|---|---|
| Audio reference on a video row | reference_audio_urls plus the prompt | Wan 3.0 at 720p: 6 x $0.125 = $0.75 |
| Audio reference on a video row | same, on Seedance 2.0 at 720p | 6 x $0.378 = $2.268 |
| Separate TTS | 300 characters of text | 300 / 1,000 x $0.0475 = $0.01425 |
| Separate TTS, then timeline render | the speech plus a silent clip | $0.01425 + $0.10 minimum render ($0.10 per output minute, minimum 1 minute) plus the clip |
Which to pick
Use an audio reference when the voice and the picture need to move together, for example a character speaking on camera. The model sees both in one pass, but you have little control over timing.
Use TTS plus a timeline when the line is fixed, the timing must be exact, or the same voice has to run across ten clips. The text is yours to edit, and a rerun of one line costs a cent or two. Omni has no audio reference, so an Omni-based pipeline needs this route anyway.
A third option is an avatar. If the speaker is a presenter rather than a character in a scene, Sume's avatar video route takes a script and a stored avatar, which keeps one voice and one face for every video. It costs $0.184 per second at the standard tier, so a 6-second clip is 6 x $0.184 = $1.104. The avatar comparison covers when that fits.
What to check before you commit
Run three clips with the same reference and the same line, and listen for drift across them. Vidu's release does not publish a limit for how many voice samples it takes or how long they may run, so read the vendor's own docs for that. On Sume, read each row's supported_input_references from GET /v1/videos/models in the video generation docs so you know if audio_url is accepted. The audio reference comparison lists the rows.
Sources
Related posts
More in Comparisons
- Waiting for Kling 4.0? Eight Sume video ids by 10-second price
Alternatives you can call today instead of waiting for Kling 4.0: eight Sume video model ids with the price of a 10-second clip, cheapest first.
- Walmart 100MB vs Amazon Sponsored Brands 500MB: bitrate each allows
Walmart caps video at 100MB, Amazon Sponsored Brands at 500MB. The average bitrate each allows for 15 to 90 seconds, plus the Sponsored Brands 1 Mbps floor.
- Wan 3.0 at 1080p ($0.25/s) is cheaper than Seedance 2.5 at 480p
Wan 3.0 at 1080p costs $0.25 a second on Sume, under Seedance 2.5 at 480p ($0.268677). So 14 s is $3.50 against $3.76, and 480p is not the budget tier there.
- Which Sume video models list auto or adaptive aspect for a start image
Seedance 2.x and Wan 3.0 list auto and adaptive aspect; Kling 3 and Omni Flash list fixed ratios only. What to send when your still is not 16:9.
Written by Sume