Vidu Q4 voice references vs Sume's reference audio inputs
Vidu Q4 takes up to 3 voice references. Sume has no Vidu row; Seedance, Wan 3.0 and MiniMax H3 accept audio references, but voice cloning is app-only, not API.

Vidu Q4 accepts up to 3 voice references, per its product page (read 2026-10-10). Sume has no Vidu row, but Seedance 2.x, Wan 3.0 and the MiniMax H3 rows accept audio references through the API; Sume's voice cloning, by contrast, is an app feature and is not available over the API.
So an audio reference on Sume is an input to a video job, not a cloned voice you can reuse.
What Vidu says
The Vidu Q4 page (read 2026-10-10) describes Image-to-Video and Reference-to-Video only, with up to 3 voice references and 1 to 15 images, 4K output and up to 16 seconds. It does not state a price, and it does not say how a voice reference is matched to a speaker, so check Vidu's own API reference before you design around it.
For this post, the key fact is the count: three voice references per request.
Which Sume rows take audio
The Video Generation docs state that Seedance 2.x, Wan 3.0, MiniMax H3 and MiniMax H3 Max accept audio and video references, while Gemini Omni Flash 1.1, Higgsfield Genjutsu and H3 Max Recast accept video references but not audio. In the request schema, non-Wan rows accept at most 3 audio files and 12 references in total; Wan 3.0 accepts up to 5 audio files.
Kling 3 and the Grok row accept no reference inputs at all.
| Sume id | Audio references | Cap |
|---|---|---|
| seedance-2.5, seedance-2 | Yes | 3 |
| wan-3.0 | Yes | 5 |
| minimax-h3, minimax-h3-max | Yes | 3 |
| gemini-omni-flash-1.1 | No | 0 |
| kling-3 | No | 0 |
| grok-imagine-video-1.5 | No | 0 |
Voice cloning is not the same thing
A voice reference in a product like Vidu is a feature of that product. Sume's own voice cloning lives in the app and cannot be called from the API. If you need a repeatable voice across many clips through code, pass the same audio file as a reference on each request to a row that accepts audio, and expect the model to treat it as a conditioning input rather than a stored voice.
Do not promise a client that the voice will match across clips. Treat each generation as independent, and review them.
- Pass
input_referenceswith an audio item on a row whose catalog entry listsaudio_urlinsupported_input_references. - Keep the audio files within the per-request caps in the table: 3 for most rows, 5 for Wan 3.0.
- Read
usage.costafter each job and compare it with the estimate before you scale up.
Cost sanity check
A 5-second Wan 3.0 clip at 720p bills $0.625 on Sume. The same duration on Seedance 2.5 at 720p bills $2.889. Neither number has a Vidu equivalent, because Vidu publishes no price on the page I read.
For a three-reference voice test, run five seconds on Wan first, listen, and only then spend on Seedance.
Designing a multi-clip voice test
Suppose you want a three-clip sequence where the same line is read in three scenes. Build the audio file once, host it on a public HTTPS URL, and attach it to each request as an audio reference. Run the first clip at the cheapest resolution, review it, then repeat for the rest.
If the first clip ignores the voice, the row may treat the audio as ambient guidance. Move to a different row in the table before you spend more. Because Sume's voice cloning is app-only, there is no API call that stores a voice and reuses it by name.
Sources
Related posts
More in Comparisons
- Voxtral Mini 4B Realtime Arabic: open weights vs a Sume STT job
Mistral's Arabic streaming model arrived Oct 8 as Apache 2.0 weights, not a service. For Arabic transcripts without hosting, Sume STT takes language_code ar.
- Wan 3.0 edit and extension at Alibaba vs Sume's Wan row
Alibaba lists Wan 3.0 editing and extension; Sume's wan-3.0 row does text, image, end frame and references only. Edit with Omni Flash 1.1 or H3 Max Recast.
- Wan 3.0 vs Seedance 2.5: the keep rate where Wan is cheaper
Wan 3.0 at 720p costs $0.63 per 5 s and Seedance 2.5 costs $2.89 on Sume. Wan wins per kept clip unless its keep rate falls below about 21.8% of Seedance's.
- Will Meta label a Sume video ad as AI? What Meta says it detects
Meta says it will detect third-party AI ads via industry-standard signals and add AI info to About this ad. The Sume docs promise nothing either way.
Written by Sume