Vidu Q4 voice references vs Sume's reference audio inputs

Vidu Q4 takes up to 3 voice references. Sume has no Vidu row; Seedance, Wan 3.0 and MiniMax H3 accept audio references, but voice cloning is app-only, not API.

4 min readSume
All posts

Vidu Q4 accepts up to 3 voice references, per its product page (read 2026-10-10). Sume has no Vidu row, but Seedance 2.x, Wan 3.0 and the MiniMax H3 rows accept audio references through the API; Sume's voice cloning, by contrast, is an app feature and is not available over the API.

So an audio reference on Sume is an input to a video job, not a cloned voice you can reuse.

What Vidu says

The Vidu Q4 page (read 2026-10-10) describes Image-to-Video and Reference-to-Video only, with up to 3 voice references and 1 to 15 images, 4K output and up to 16 seconds. It does not state a price, and it does not say how a voice reference is matched to a speaker, so check Vidu's own API reference before you design around it.

For this post, the key fact is the count: three voice references per request.

Which Sume rows take audio

The Video Generation docs state that Seedance 2.x, Wan 3.0, MiniMax H3 and MiniMax H3 Max accept audio and video references, while Gemini Omni Flash 1.1, Higgsfield Genjutsu and H3 Max Recast accept video references but not audio. In the request schema, non-Wan rows accept at most 3 audio files and 12 references in total; Wan 3.0 accepts up to 5 audio files.

Kling 3 and the Grok row accept no reference inputs at all.

Audio reference support on Sume (checked 2026-10-10)
Sume idAudio referencesCap
seedance-2.5, seedance-2Yes3
wan-3.0Yes5
minimax-h3, minimax-h3-maxYes3
gemini-omni-flash-1.1No0
kling-3No0
grok-imagine-video-1.5No0

Voice cloning is not the same thing

A voice reference in a product like Vidu is a feature of that product. Sume's own voice cloning lives in the app and cannot be called from the API. If you need a repeatable voice across many clips through code, pass the same audio file as a reference on each request to a row that accepts audio, and expect the model to treat it as a conditioning input rather than a stored voice.

Do not promise a client that the voice will match across clips. Treat each generation as independent, and review them.

  • Pass input_references with an audio item on a row whose catalog entry lists audio_url in supported_input_references.
  • Keep the audio files within the per-request caps in the table: 3 for most rows, 5 for Wan 3.0.
  • Read usage.cost after each job and compare it with the estimate before you scale up.

Cost sanity check

A 5-second Wan 3.0 clip at 720p bills $0.625 on Sume. The same duration on Seedance 2.5 at 720p bills $2.889. Neither number has a Vidu equivalent, because Vidu publishes no price on the page I read.

For a three-reference voice test, run five seconds on Wan first, listen, and only then spend on Seedance.

Designing a multi-clip voice test

Suppose you want a three-clip sequence where the same line is read in three scenes. Build the audio file once, host it on a public HTTPS URL, and attach it to each request as an audio reference. Run the first clip at the cheapest resolution, review it, then repeat for the rest.

If the first clip ignores the voice, the row may treat the audio as ambient guidance. Move to a different row in the table before you spend more. Because Sume's voice cloning is app-only, there is no API call that stores a voice and reuses it by name.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume