PixVerse V6 adds native audio: which Sume video ids make sound
PixVerse V6 ships camera work and native audio. The Sume video docs name no PixVerse id, so here is how to find models that generate audio and read the flag.

The PixVerse blog describes V6 as adding camera work and native audio. The Sume video docs read for this post name no PixVerse model id, so the practical question for a Sume user is different: which catalog models generate audio, and how do you check?
What the PixVerse page says
The PixVerse blog also describes R2, a real-time world model, and an extended Series C bringing total funding to $439M. Those are company facts, not API facts, and none of them changes what the Sume catalog lists.
Audio on Sume
Every row in GET /v1/videos/models carries a generate_audio boolean that says whether the model can produce an audio track. The docs name several audio-capable models in prose.
| Model | What the docs say about audio |
|---|---|
seedance-2 | Example catalog row shows generate_audio: true |
minimax-h3-max | Native stereo audio |
gemini-omni-flash-1.1 | Native synced audio |
sume/auto | The current Auto family always generates audio |
Reading the flag in code
The generate_audio request field defaults to the model's audio capability. For sume/auto the legacy Video 1.0 docs say to omit the field, since the Auto family always generates audio.
import os
import requests
resp = requests.get(
"https://api.sume.com/v1/videos/models",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
timeout=30,
)
resp.raise_for_status()
with_audio = [m["id"] for m in resp.json()["data"] if m.get("generate_audio")]
print(with_audio)Audio and captions
Native audio changes the shape of a pipeline. With a silent model you need a separate voice and music step, then a mix. With native audio the clip arrives with sound, which saves a step but gives you less control over levels and wording.
If a script needs exact words on screen, add captions after generation rather than trusting the model to render text or to speak the exact line. The Sume docs list a caption workflow separately from video generation.
- Need exact spoken words: generate a voice track separately.
- Need only ambience: native audio is usually enough.
- Need captions: add them after generation.
- Always play the clip with sound on before sign-off.
What to do next
Run the snippet against the live catalog. If you need a model with a specific feel, such as the camera movement PixVerse markets for V6, test short clips from the audio-capable ids rather than trusting a feature list. The video docs describe every field you can filter on.
Sources
Related posts
More in Models
- Repeated image edits without drift: a native-resolution test plan
Recraft says Ideogram 4.5 is built for stable repeated edits at native resolution without downsampling. A five-pass test to check any edit model for drift.
- Scribe v2, Realtime or Medical: which model for interview clips
ElevenLabs lists Scribe v2 with up to 32 speakers and 90+ languages, Realtime at about 150 ms, and Medical with 35% fewer errors. Which fits interview clips.
- Seedance 2.5 on BytePlus excludes the US: check access first
BytePlus lists Seedance 2.5 on ModelArk in supported markets excluding the United States. What that means for a US team, and how to read the Sume catalog.
- Self-host MiniMax H3 or use an API: a decision table
MiniMax released H3 as an open-weight omni-modal model on Jul 31, 2026. A decision table for running it yourself versus calling a hosted video API.
Written by Sume