Grok Imagine's 3 voice references vs Sume reference-audio rows
xAI's Grok Imagine 1.5 takes up to 3 voice references. Sume's Grok row takes none; Seedance, Wan 3.0 and MiniMax accept reference audio under limits.

Sume's Grok Imagine row accepts no audio at all, so the 3 voice references that xAI documents for grok-imagine-video-1.5 have no equivalent there. For reference audio on Sume, use Seedance, Wan 3.0, MiniMax H3 or MiniMax H3 Max, each with its own limits (xAI video guide, read 2026-10-10).
The Sume Grok row also rejects reference_audio_urls and generate_audio, so the failure is a clear 400 and not a silent drop.
What each side documents
xAI's guide lists up to 14 reference images and up to 3 voice references on the 1.5 model. It does not say here how long a voice reference may be, so no duration claim is made for it.
Sume's catalog flags reference_audios per row. The Video Router schema adds that reference_audio_urls needs at least one reference image or reference video alongside it, so audio alone is not a valid reference set.
| Row | Reference audio | Limit noted in the catalog |
|---|---|---|
| xAI grok-imagine-video-1.5 | Up to 3 voice references | Not stated in the page text |
| Sume Grok Imagine Video 1.5 | No | Rejected |
| Seedance 2.5 and 2.0 family | Yes | Audio and video references accepted |
| Wan 3.0 | Yes | At most 5, 15 s total |
| MiniMax H3 and H3 Max | Yes | At most 3, each 2 to 15 s, 15 s combined; audio cannot be the only reference |
| Gemini Omni Flash 1.1 | No | Rejected |
| Kling Video v3 Pro | No | No reference inputs |
Voice reference vs audio reference
A voice reference in xAI's wording steers how the speaker sounds. Sume's reference_audio_urls on Seedance, Wan or MiniMax is an input clip the model can use as guidance. Do not assume the two behave the same way; test a short clip first.
Voice cloning on Sume is an app feature, not an API one. If you need a particular voice on the soundtrack through the API, the safer route is to generate the picture with sound on a row that makes native audio, then replace or mix the audio track with a separate step.
A port plan
For a Grok job that used voice references:
- Pick Wan 3.0 (2 to 30 s, up to 5 audio refs) when you also have reference images.
- Pick MiniMax H3 (5 to 15 s) when your clips are at least 5 seconds and the audio clips are 2 to 15 s each.
- Keep the reference set to at least one image or video plus the audio.
- Check
supported_input_referenceson the row foraudio_urlbefore you submit.
Sources
Related posts
More in Comparisons
- Grok Imagine Lite draft plus Sume upscale: what a 10 s clip totals
xAI lists Lite at $0.020 per second and 1.5 at $0.080. Add Sume video upscale at $0.009 per input second and a 10 s clip totals $0.29. Read 2026-10-10.
- Grok Imagine video edit on xAI vs the video_url error on Sume
xAI edits video with grok-imagine-video; Sume accepts video_url only on Gemini Omni Flash 1.1. The exact error, the edit limits and what to send instead.
- Hedra batch_size 1-8 variations vs Sume's one job per request
Hedra's API returns 1 to 8 variations per request; Sume avatar videos are one job each. How to get variations on Sume with previews and keys (read 2026-10-10).
- Hedra Character 3 vs Sume Avatar 1.0: per-second cost, 30 and 60 s
Hedra lists Character 3 at 2.5 to 6.25 cents per second; Sume Avatar 1.0 is $0.184 to $0.55 per second at 720p. Totals at 30 and 60 s (read 2026-10-10).
Written by Sume