Kling 3 lip sync with your own audio: what each API does
Kling lists lip sync driven by text or audio. Sume's kling-3 takes no audio input, only a generate_audio switch; models with audio references are listed.

Kling lists Lip Sync as a capability of all its model versions, combined with text or audio to drive a character's mouth, but Sume's kling-3 takes no audio file: it accepts no input_references, only a generate_audio on/off switch for sound the model makes itself. To drive a video from your own audio on Sume, use a model that lists audio_url references, or a tool built for talking video.
Kling's capability is from its capability map; Sume's limits from Video generation and code, read 2026-09-29.
What does Kling say about lip sync and voice?
The map's Global Capabilities table lists Lip Sync for all model versions, "combined with text or audio to drive the mouth shape of characters in the video", and Text to Audio and Video to Audio for all versions. For Kling 3.0 text-to-video it marks Native Audio supported and Voice Control (human voice) not supported.
What can I send on Sume?
The docs say audio and video references are honored by the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max. "Honored" means the reference is used as guidance; the docs do not promise frame-accurate mouth sync, so test a short clip.
| Model | Audio behavior |
|---|---|
kling-3 | generate_audio on or off; no audio input |
seedance-2.5, seedance-2, seedance-2-fast, seedance-2-mini | Audio references honored; optional generate_audio |
wan-3.0 | Audio references honored |
minimax-h3, minimax-h3-max | Audio references honored |
gemini-omni-flash-1.1 | Native audio; no audio references |
grok-imagine-video-1.5 | No audio |
How do I send an audio reference?
Add an audio_url entry to input_references, alongside an image reference.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: audio-ref-001" \
-d '{
"model": "seedance-2.5",
"prompt": "The woman in the photo speaks to camera, warm light.",
"input_references": [
{"type":"image_url","image_url":{"url":"https://example.com/face.png"}},
{"type":"audio_url","audio_url":{"url":"https://example.com/voice.mp3"}}
],
"duration": 8
}'Is there a better tool for talking video?
For a face that must speak your exact audio, the tools built for it are covered in talking video: avatar vs lip sync vs motion control and lip sync from a photo and audio.
Sources
Related posts
More in Models
- Kling 3 native audio: what the on and off price difference is
On Sume, kling-3 costs $0.14 a second silent and $0.21 with audio. Audio is on unless you send generate_audio false, so the default is the higher price.
- Lyria 3 Pro or Lyria 3.5: which model id do I send?
Sume's Music Router accepts sume/music-auto, lyria-3.5 and lyria-3-pro. What each id does, how to pin one, and what Google lists for the models.
- Lyria 3.5 API: model id, request and how to call it
Lyria 3.5 is in the Gemini API as lyria-3.5. Here is what Google's page says about it and how to call it from Sume with one music request.
- AI music generator with vocals: Lyria 3.5 lyrics by API
Lyria 3.5 can sing. How to ask for vocals or an instrumental, steer lyrics with section tags, and read the lyrics back from a Sume music job.
Written by Sume