Kling API avatar: one image plus audio, 2 to 300 seconds
Kling's API has an Avatar endpoint: one reference image plus a 2–300 second audio track becomes a talking video, billed per second. Inputs and Sume options.
Yes, Kling's API has an avatar endpoint: POST /v1/videos/avatar/image2video animates one reference image into a talking video driven by an audio track of 2 to 300 seconds, either a sound file or audio from Kling's own TTS. It runs in std or pro mode and is billed per second.
Kling's facts come from its Avatar, Video pricing and Authentication API pages, read on 2026-09-28. Sume does not route Kling's Avatar model; the Sume options further down come from the Sume API reference and the Generate avatar video docs.
What does the Kling Avatar API take?
The request is JSON with a Bearer API key from the Kling AI console. The result is a task you query or get a callback_url notification for.
| Item | Kling Avatar API |
|---|---|
| Endpoint | POST /v1/videos/avatar/image2video on https://api-singapore.klingai.com |
image (required) | URL or raw Base64; .jpg, .jpeg or .png; ≤ 10MB; ≥ 300px; aspect ratio 1:2.5 to 2.5:1 |
sound_file or audio_id | One of the two. File: .mp3, .wav, .m4a or .aac, max 5MB, 2–300 seconds. audio_id: Kling TTS audio from the last 30 days |
prompt | Optional; avatar actions, emotions and camera movements; 2500 characters max |
mode | std (default) or pro |
| Task status | submitted, processing, succeed, failed |
| Price | 0.4 Units ($0.056) per second in the 720P column, 0.8 Units ($0.112) per second in the 1080P column; Avatar TTS 0.05 Units ($0.007) per call |
How long do Kling avatar videos stay available?
Kling's response notes that generated videos "will be cleared after 30 days", so copy the file from task_result.videos[].url once the task reaches succeed. Kling's pricing table does not say which mode maps to which price column; the task result reports the units deducted in final_unit_deduction.
Can I make the same image-plus-audio video on Sume?
Yes, with a different model: POST /v1/veed/fabric-1.0 (VEED Fabric 1.0) turns a still image, or a ready avatar's still, plus audio into a talking clip. It costs $0.1875 per audio second (720p) on API pricing, and Lip sync API walks through the request. Where its inputs differ from Kling's Avatar call:
- Picture: exactly one visual source, a public HTTPS
image_urlor a ready avatar'savatar_id/avatar_handle. - Audio: Kling takes a
sound_fileor a Kling TTSaudio_id. Sume'saudio_urlmust be a file on Sume's media host, typically a TTS segment, of at most 10 MB; other hosts are refused. - Length and size: the audio's length goes in
duration_seconds(1 to 300), andresolutionis480por720p(default720p).
What if I only have a script, not audio?
Kling's route takes TTS audio from its own TTS endpoint. On Sume, the script-in route is Avatar 1.0, POST /v1/avatar-1.0/talking-video, which turns a ready avatar and a script into a talking video.
- Scripts are accepted when Sume estimates the video at 4 to 60 seconds; split longer scripts into several jobs.
aspect_ratiois1:1,3:4,9:16(default),4:3or16:9, and 720p is the documented resolution.- In current code the avatar speaks English only. For other languages, generate speech with TTS 1.0 (
POST /v1/tts-1.0/generate, withlanguageset) and lip-sync it with Fabric.
Sources
Related posts
More in Sume Avatar 1.0
- Difference between lip sync and dubbing: which do you need?
Dubbing replaces the speech in a video; lip sync matches a mouth to the audio. Lip-sync dubbing does both. What each means, and which one you need.
- How long should a microlearning video be? Length and AI
No fixed standard: one objective per video, and research on course videos favors 6 minutes or less. How to size lessons and make them with an AI avatar.
- Real-time AI avatar vs video avatar API: the difference
A real-time AI avatar talks live in a session; a video avatar API renders a finished file from a script. Which vendors sell which, and where Sume fits.
- Runway Characters API: avatars, live sessions and videos
Runway Characters builds avatars from one image, a voice and a personality, for live sessions or rendered videos. Endpoints, limits and pricing.
Written by Sume