Grok Imagine video sound: xAI audio vs Sume's silent row
xAI's Grok Imagine video makes audio unless generate_audio is False. Sume's grok-imagine-video-1.5 row lists audio: false and rejects generate_audio.

Grok Imagine video sound is an xAI-side feature: its docs turn audio generation on and let you disable it with generate_audio=False. The Sume Grok row lists audio: false and rejects generate_audio, so it returns silent video.
xAI's facts are from its video docs and July 31, 2026 post, read 2026-09-30. Sume's are from its catalog and Video Router docs.
What does xAI offer for sound?
Its docs list audio generation, disabled with generate_audio=False, and up to 3 voices per request for reference audio. The July 31 post adds voice consistency through character images and voice references, with voice reference support in the API on request.
What does the Sume row do?
The catalog sets audio: false and reference_audios: false, and its constraints include "no aspect_ratio / bitrate_mode / generate_audio". Treat the output as silent and add sound separately.
| Where | Audio |
|---|---|
| xAI, Grok | Generated; off with generate_audio=False |
| Sume, Grok | None (audio: false) |
Sume, gemini-omni-flash-1.1 | Native synced audio always on |
Sume, wan-3.0 | audio: true in the catalog |
Which Sume ids carry sound?
The docs say gemini-omni-flash-1.1 has native synced audio always on, and generate_audio: false is rejected there. MiniMax H3 Max is documented with native stereo audio. For a general walkthrough see AI video with sound.
What should I do for a Grok clip that needs sound?
Call xAI directly, or keep the silent clip and add music or a voiceover afterwards with a separate tool. Read audio in the catalog before assuming a track exists.
Sources
Related posts
More in Models
- Hy-Image-3.5-Preview API limits vs Sume's image limits
Tencent's Hy-Image-3.5-Preview takes 256 to 8192 px edges, up to 4096x4096 and 20 references. Sume lists no Hy Image model; its own limits are below.
- Ideogram 4.5 Precise Edit API: mask and reference limits
Ideogram 4.5 Precise Edit takes 4 reference images, 3 when a mask is sent. Sume lists Ideogram V3, and mask_url there is only for ChatGPT Image 2.5.
- How many reference images? 10 on Image 1.0, 16 on GPT Image 2.5
Image 1.0 image_urls takes 1 to 10 public HTTPS URLs; ChatGPT Image 2.5 takes up to 16 references. Text-only models reject references.
- Image quality default on Sume: low on Image 1.0, high on GPT 2.5
Image 1.0 defaults quality to low; ChatGPT Image 2.5 on POST /v1/images defaults to high when omitted. Values, and when to raise quality.
Written by Sume