Prism audio tags vs Sume generate_audio and audio references
Prism's README describes music, sfx and speech tags. Sume gives a generate_audio flag and audio references on some ids. Which ids, and the limits.

Prism's README describes structured audio tags, <music>, <sfx> and <speech>, plus a separate audio caption field in its dataset format. Sume has no such tag syntax in the docs this page read. What you control on Sume is a generate_audio flag on ids that allow it, and reference audio files on a few ids.
The Prism details are from its repository, read on 2026-10-11. The README excerpt I read describes them in the context of training data, so check the inference scripts before assuming you can type tags into a prompt.
What Sume exposes for sound
Sume's video generation docs define generate_audio as telling the model to produce an audio track or not, with the default set by the model's audio capability. The catalog code shows which ids have audio, and the docs list which ones accept audio references.
| Sume id | Makes audio | Audio reference input |
|---|---|---|
| seedance-2.5, seedance-2, -fast, -mini | yes | yes |
| wan-3.0 | yes | yes, up to 5 files |
| minimax-h3, minimax-h3-max | yes (native stereo) | yes |
| kling-3 | yes, priced higher | no reference inputs |
| gemini-omni-flash-1.1 | always on | no (image and video only) |
| grok-imagine-video-1.5 | no | no |
What sound costs on kling-3
kling-3 is the one row where the catalog prices audio separately: $0.14 per second without audio and $0.21 with it, both at 1080p. A 10-second clip is $1.40 silent and $2.10 with sound, so the audio adds $0.70, or 50 percent. The other per-second rows in the catalog code carry no separate audio rate, and the Seedance family prices by video tokens instead.
How audio references behave
On the Video Router request schema, a reference audio file must come with at least one reference image or video, and wan-3.0 accepts at most 5 reference_audio_urls. The video docs list the Seedance 2.x ids, Wan 3.0, MiniMax H3 and MiniMax H3 Max as accepting audio and video references, and say Gemini Omni Flash 1.1 and the Higgsfield and Recast ids accept video but not audio.
That makes an audio reference a style or content guide supplied with pictures, not a standalone soundtrack upload. If you need a track laid exactly under a finished clip, use Sume's Timeline audio tools (see the models overview) rather than a generation reference.
Choosing for sound control
- Want separate control of music, effects and speech: Prism's tags point that way, but verify them for inference first.
- Want to feed a voice or a beat into the generation: pick an id from the table that accepts audio references, and send an image or video with it.
- Want sound you cannot turn off:
gemini-omni-flash-1.1always carries audio. - Want a silent clip:
generate_audioset to false where the id allows it, orgrok-imagine-video-1.5, which has none.
Sources
Related posts
More in Comparisons
- Stability's Stable Image edit tools, matched to Sume image routes
Which of Stability's Stable Image edit, upscale and control tools have a Sume route: mask edits, background removal and upscale yes, outpaint and sketch no.
- Submagic caps video at 2, 5 or 30 min; Sume prices up to 60 seconds
Submagic allows 2, 5 or 30 minute videos by plan. Sume's $0.20 caption job is priced for videos up to 60 seconds. Which one fits which video length.
- Synthesia Syren writes video as code: what Sume has instead
Syren is Synthesia's prompt-to-video agent in early access. Sume has no equivalent; it has a script-to-avatar API and a timeline. The gap, mapped.
- Syren mixes the first 24 audio tracks; Sume's timeline takes 20 parts
Syren mixes only the first 24 audio tracks in an export. Sume's timeline takes one spine in up to 20 gapless parts plus one bed. What fits where.
Written by Sume