Kling 4.0 voice reference vs Sume reference_audio_urls
Kling 4.0 Omni Reference accepts voice references. Sume takes 1-3 reference_audio_urls on models such as MiniMax H3; here is what each source documents.

Kling 4.0's Omni Reference supports voice references, and on Sume the closest control is reference_audio_urls: one to three audio URLs, sent with at least one reference image or video, on models that honor audio references. Sume's docs describe the field as an audio reference; they do not promise that a voice carries over, so test it.
What each side documents
The Kling column is from the Kling 4.0 announcement. The Sume column is from the video generation and Video 1.0 docs.
| Item | Kling 4.0 | Sume |
|---|---|---|
| Voice reference | Omni Reference supports voice references | reference_audio_urls, 1-3 audio URLs |
| Needs another reference | Not stated in what I read | At least one reference image or video |
| Editing inputs | Up to five video inputs, 30 s combined duration | Gemini Omni Flash 1.1 video edit takes one video_url |
| Where audio references work | Not stated | Seedance 2.x, Wan 3.0, MiniMax H3, MiniMax H3 Max |
Send audio references on Sume
The Video Router route takes flat fields. The call below sends one image and one audio reference to minimax-h3, which accepts 5 to 15 seconds at native 480p or 768p. Models that list audio_url in supported_input_references accept the audio; Gemini Omni Flash 1.1 and the Genjutsu and Recast models accept video references but not audio.
On the newer route, POST /v1/videos, the same idea is an input_references array. Check supported_input_references on GET /v1/videos/models for the model you pick.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: voice-ref-001" \
-d '{
"model": "minimax-h3",
"prompt": "The presenter greets the viewer and holds up the product",
"reference_image_urls": ["https://example.com/presenter.png"],
"reference_audio_urls": ["https://example.com/voice-sample.mp3"],
"resolution": "768p",
"duration": 8,
"mode": "async"
}'How to test whether the voice matched
Render the same line with and without the audio reference, then listen back to back. Run video inspect with transcribe set to true to confirm the words are right. Matching a voice is a listening judgment, and a transcript only checks the words.
If you need an exact, repeatable voice across many lines, a separate text-to-speech step is a different design from a reference input; this post only covers the reference route.
Sources
Related posts
More in Comparisons
- Kling 4.0 vs Kling 3.0: the spec differences in one dated table
Length, resolution, HDR, references, keyframes, audio and prompt size, Kling 4.0 against 3.0 as stated on Kling's pages, plus how each maps to a Sume request.
- Kling V3 Pro vs Standard vs O3 per minute
Hedra's Artificial Analysis table puts Kling V3 Pro at $20.16 per minute and V3 Standard at $15.12. How to compare per-minute prices with Sume per-second SKUs.
- Live avatar agent or rendered video: D-ID vs Sume
D-ID V4 Expressive Visual Agents are live, LLM-connected avatars. Sume avatar video is rendered from a script in 4 to 60 s. Which one fits which job.
- LTX-2.5 license ($10M ARR line) vs hosted video ids
LTX-2.x Community license is free under $10M ARR and paid above it. Here is what changes if you call a hosted Sume video id instead of running weights.
Written by Sume