Grok Imagine's 3 voice references vs Sume reference-audio rows

xAI's Grok Imagine 1.5 takes up to 3 voice references. Sume's Grok row takes none; Seedance, Wan 3.0 and MiniMax accept reference audio under limits.

4 min readSume
All posts

Sume's Grok Imagine row accepts no audio at all, so the 3 voice references that xAI documents for grok-imagine-video-1.5 have no equivalent there. For reference audio on Sume, use Seedance, Wan 3.0, MiniMax H3 or MiniMax H3 Max, each with its own limits (xAI video guide, read 2026-10-10).

The Sume Grok row also rejects reference_audio_urls and generate_audio, so the failure is a clear 400 and not a silent drop.

What each side documents

xAI's guide lists up to 14 reference images and up to 3 voice references on the 1.5 model. It does not say here how long a voice reference may be, so no duration claim is made for it.

Sume's catalog flags reference_audios per row. The Video Router schema adds that reference_audio_urls needs at least one reference image or reference video alongside it, so audio alone is not a valid reference set.

Reference audio: xAI guide (read 2026-10-10) and Sume catalog in the repo on 2026-10-10
RowReference audioLimit noted in the catalog
xAI grok-imagine-video-1.5Up to 3 voice referencesNot stated in the page text
Sume Grok Imagine Video 1.5NoRejected
Seedance 2.5 and 2.0 familyYesAudio and video references accepted
Wan 3.0YesAt most 5, 15 s total
MiniMax H3 and H3 MaxYesAt most 3, each 2 to 15 s, 15 s combined; audio cannot be the only reference
Gemini Omni Flash 1.1NoRejected
Kling Video v3 ProNoNo reference inputs

Voice reference vs audio reference

A voice reference in xAI's wording steers how the speaker sounds. Sume's reference_audio_urls on Seedance, Wan or MiniMax is an input clip the model can use as guidance. Do not assume the two behave the same way; test a short clip first.

Voice cloning on Sume is an app feature, not an API one. If you need a particular voice on the soundtrack through the API, the safer route is to generate the picture with sound on a row that makes native audio, then replace or mix the audio track with a separate step.

A port plan

For a Grok job that used voice references:

  • Pick Wan 3.0 (2 to 30 s, up to 5 audio refs) when you also have reference images.
  • Pick MiniMax H3 (5 to 15 s) when your clips are at least 5 seconds and the audio clips are 2 to 15 s each.
  • Keep the reference set to at least one image or video plus the audio.
  • Check supported_input_references on the row for audio_url before you submit.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume