PixVerse V6 adds native audio: which Sume video ids make sound

PixVerse V6 ships camera work and native audio. The Sume video docs name no PixVerse id, so here is how to find models that generate audio and read the flag.

4 min readSume
All posts

The PixVerse blog describes V6 as adding camera work and native audio. The Sume video docs read for this post name no PixVerse model id, so the practical question for a Sume user is different: which catalog models generate audio, and how do you check?

What the PixVerse page says

The PixVerse blog also describes R2, a real-time world model, and an extended Series C bringing total funding to $439M. Those are company facts, not API facts, and none of them changes what the Sume catalog lists.

Audio on Sume

Every row in GET /v1/videos/models carries a generate_audio boolean that says whether the model can produce an audio track. The docs name several audio-capable models in prose.

Audio statements in the Sume video docs
ModelWhat the docs say about audio
seedance-2Example catalog row shows generate_audio: true
minimax-h3-maxNative stereo audio
gemini-omni-flash-1.1Native synced audio
sume/autoThe current Auto family always generates audio

Reading the flag in code

The generate_audio request field defaults to the model's audio capability. For sume/auto the legacy Video 1.0 docs say to omit the field, since the Auto family always generates audio.

import os
import requests

resp = requests.get(
    "https://api.sume.com/v1/videos/models",
    headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
    timeout=30,
)
resp.raise_for_status()
with_audio = [m["id"] for m in resp.json()["data"] if m.get("generate_audio")]
print(with_audio)

Audio and captions

Native audio changes the shape of a pipeline. With a silent model you need a separate voice and music step, then a mix. With native audio the clip arrives with sound, which saves a step but gives you less control over levels and wording.

If a script needs exact words on screen, add captions after generation rather than trusting the model to render text or to speak the exact line. The Sume docs list a caption workflow separately from video generation.

  • Need exact spoken words: generate a voice track separately.
  • Need only ambience: native audio is usually enough.
  • Need captions: add them after generation.
  • Always play the clip with sound on before sign-off.

What to do next

Run the snippet against the live catalog. If you need a model with a specific feel, such as the camera movement PixVerse markets for V6, test short clips from the audio-capable ids rather than trusting a feature list. The video docs describe every field you can filter on.

Sources

Related posts

More in Models

All Models posts

Written by Sume