Kling 4.0 voice reference vs Sume reference_audio_urls

Kling 4.0 Omni Reference accepts voice references. Sume takes 1-3 reference_audio_urls on models such as MiniMax H3; here is what each source documents.

4 min readSume
All posts

Kling 4.0's Omni Reference supports voice references, and on Sume the closest control is reference_audio_urls: one to three audio URLs, sent with at least one reference image or video, on models that honor audio references. Sume's docs describe the field as an audio reference; they do not promise that a voice carries over, so test it.

What each side documents

The Kling column is from the Kling 4.0 announcement. The Sume column is from the video generation and Video 1.0 docs.

Voice and reference inputs (read 2026-10-03)
ItemKling 4.0Sume
Voice referenceOmni Reference supports voice referencesreference_audio_urls, 1-3 audio URLs
Needs another referenceNot stated in what I readAt least one reference image or video
Editing inputsUp to five video inputs, 30 s combined durationGemini Omni Flash 1.1 video edit takes one video_url
Where audio references workNot statedSeedance 2.x, Wan 3.0, MiniMax H3, MiniMax H3 Max

Send audio references on Sume

The Video Router route takes flat fields. The call below sends one image and one audio reference to minimax-h3, which accepts 5 to 15 seconds at native 480p or 768p. Models that list audio_url in supported_input_references accept the audio; Gemini Omni Flash 1.1 and the Genjutsu and Recast models accept video references but not audio.

On the newer route, POST /v1/videos, the same idea is an input_references array. Check supported_input_references on GET /v1/videos/models for the model you pick.

curl -X POST https://api.sume.com/v1/video-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: voice-ref-001" \
  -d '{
    "model": "minimax-h3",
    "prompt": "The presenter greets the viewer and holds up the product",
    "reference_image_urls": ["https://example.com/presenter.png"],
    "reference_audio_urls": ["https://example.com/voice-sample.mp3"],
    "resolution": "768p",
    "duration": 8,
    "mode": "async"
  }'

How to test whether the voice matched

Render the same line with and without the audio reference, then listen back to back. Run video inspect with transcribe set to true to confirm the words are right. Matching a voice is a listening judgment, and a transcript only checks the words.

If you need an exact, repeatable voice across many lines, a separate text-to-speech step is a different design from a reference input; this post only covers the reference route.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume