Which Sume video models accept an audio reference?

Seven Sume video models take reference audio: the four Seedance 2.x ids, Wan 3.0 and the two MiniMax H3 models. Kling, Grok and Gemini Omni Flash do not.

4 min readSume
All posts

Seven ids in Sume's video catalog accept reference audio: seedance-2.5, seedance-2, seedance-2-fast, seedance-2-mini, wan-3.0, minimax-h3 and minimax-h3-max. The docs group the four Seedance ids as "the Seedance 2.x models". kling-3, grok-imagine-video-1.5, gemini-omni-flash-1.1 and the Higgsfield motion-transfer model do not.

Source: the capability flags in Sume's catalog and the video docs, read 2026-10-01. A model that does not list audio_url in supported_input_references rejects it.

What are the per-model limits?

Only Wan and MiniMax publish numbers in the catalog. Seedance limits are not stated there, so check the live model list before you build around a count.

Audio-reference limits in Sume's catalog, read 2026-10-01.
ModelReference audio
wan-3.0up to 5 files, 15 seconds combined
minimax-h3, minimax-h3-maxup to 3 files of 2 to 15 seconds, 15 combined; images, videos and audios up to 12 together; audio cannot be the only reference
seedance-2.5, seedance-2, seedance-2-fast, seedance-2-minisupported; limits not listed in the catalog
kling-3, grok-imagine-video-1.5, gemini-omni-flash-1.1not accepted

Which models reject reference audio?

Per the catalog, kling-3, grok-imagine-video-1.5, gemini-omni-flash-1.1 and the Higgsfield motion-transfer model do not list audio_url, so they reject it.

How do I check a model before I send audio?

Read supported_input_references on the model list; it should include audio_url. Then send the file in input_references.

curl https://api.sume.com/v1/videos/models \
  -H "Authorization: Bearer $SUME_API_KEY" \
  | grep -o '"supported_input_references":[^]]*]'

Sources

Related posts

More in Models

All Models posts

Written by Sume