Prism audio tags vs Sume generate_audio and audio references

Prism's README describes music, sfx and speech tags. Sume gives a generate_audio flag and audio references on some ids. Which ids, and the limits.

5 min readSume
All posts

Prism's README describes structured audio tags, <music>, <sfx> and <speech>, plus a separate audio caption field in its dataset format. Sume has no such tag syntax in the docs this page read. What you control on Sume is a generate_audio flag on ids that allow it, and reference audio files on a few ids.

The Prism details are from its repository, read on 2026-10-11. The README excerpt I read describes them in the context of training data, so check the inference scripts before assuming you can type tags into a prompt.

What Sume exposes for sound

Sume's video generation docs define generate_audio as telling the model to produce an audio track or not, with the default set by the model's audio capability. The catalog code shows which ids have audio, and the docs list which ones accept audio references.

Audio controls by Sume video id (docs and catalog code, read 2026-10-11)
Sume idMakes audioAudio reference input
seedance-2.5, seedance-2, -fast, -miniyesyes
wan-3.0yesyes, up to 5 files
minimax-h3, minimax-h3-maxyes (native stereo)yes
kling-3yes, priced higherno reference inputs
gemini-omni-flash-1.1always onno (image and video only)
grok-imagine-video-1.5nono

What sound costs on kling-3

kling-3 is the one row where the catalog prices audio separately: $0.14 per second without audio and $0.21 with it, both at 1080p. A 10-second clip is $1.40 silent and $2.10 with sound, so the audio adds $0.70, or 50 percent. The other per-second rows in the catalog code carry no separate audio rate, and the Seedance family prices by video tokens instead.

How audio references behave

On the Video Router request schema, a reference audio file must come with at least one reference image or video, and wan-3.0 accepts at most 5 reference_audio_urls. The video docs list the Seedance 2.x ids, Wan 3.0, MiniMax H3 and MiniMax H3 Max as accepting audio and video references, and say Gemini Omni Flash 1.1 and the Higgsfield and Recast ids accept video but not audio.

That makes an audio reference a style or content guide supplied with pictures, not a standalone soundtrack upload. If you need a track laid exactly under a finished clip, use Sume's Timeline audio tools (see the models overview) rather than a generation reference.

Choosing for sound control

  • Want separate control of music, effects and speech: Prism's tags point that way, but verify them for inference first.
  • Want to feed a voice or a beat into the generation: pick an id from the table that accepts audio references, and send an image or video with it.
  • Want sound you cannot turn off: gemini-omni-flash-1.1 always carries audio.
  • Want a silent clip: generate_audio set to false where the id allows it, or grok-imagine-video-1.5, which has none.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume