Kling API avatar: one image plus audio, 2 to 300 seconds

Kling's API has an Avatar endpoint: one reference image plus a 2–300 second audio track becomes a talking video, billed per second. Inputs and Sume options.

5 min readSume
All posts

Yes, Kling's API has an avatar endpoint: POST /v1/videos/avatar/image2video animates one reference image into a talking video driven by an audio track of 2 to 300 seconds, either a sound file or audio from Kling's own TTS. It runs in std or pro mode and is billed per second.

Kling's facts come from its Avatar, Video pricing and Authentication API pages, read on 2026-09-28. Sume does not route Kling's Avatar model; the Sume options further down come from the Sume API reference and the Generate avatar video docs.

What does the Kling Avatar API take?

The request is JSON with a Bearer API key from the Kling AI console. The result is a task you query or get a callback_url notification for.

From Kling's Avatar, Video pricing and Authentication pages, read 2026-09-28.
ItemKling Avatar API
EndpointPOST /v1/videos/avatar/image2video on https://api-singapore.klingai.com
image (required)URL or raw Base64; .jpg, .jpeg or .png; ≤ 10MB; ≥ 300px; aspect ratio 1:2.5 to 2.5:1
sound_file or audio_idOne of the two. File: .mp3, .wav, .m4a or .aac, max 5MB, 2–300 seconds. audio_id: Kling TTS audio from the last 30 days
promptOptional; avatar actions, emotions and camera movements; 2500 characters max
modestd (default) or pro
Task statussubmitted, processing, succeed, failed
Price0.4 Units ($0.056) per second in the 720P column, 0.8 Units ($0.112) per second in the 1080P column; Avatar TTS 0.05 Units ($0.007) per call

How long do Kling avatar videos stay available?

Kling's response notes that generated videos "will be cleared after 30 days", so copy the file from task_result.videos[].url once the task reaches succeed. Kling's pricing table does not say which mode maps to which price column; the task result reports the units deducted in final_unit_deduction.

Can I make the same image-plus-audio video on Sume?

Yes, with a different model: POST /v1/veed/fabric-1.0 (VEED Fabric 1.0) turns a still image, or a ready avatar's still, plus audio into a talking clip. It costs $0.1875 per audio second (720p) on API pricing, and Lip sync API walks through the request. Where its inputs differ from Kling's Avatar call:

  • Picture: exactly one visual source, a public HTTPS image_url or a ready avatar's avatar_id / avatar_handle.
  • Audio: Kling takes a sound_file or a Kling TTS audio_id. Sume's audio_url must be a file on Sume's media host, typically a TTS segment, of at most 10 MB; other hosts are refused.
  • Length and size: the audio's length goes in duration_seconds (1 to 300), and resolution is 480p or 720p (default 720p).

What if I only have a script, not audio?

Kling's route takes TTS audio from its own TTS endpoint. On Sume, the script-in route is Avatar 1.0, POST /v1/avatar-1.0/talking-video, which turns a ready avatar and a script into a talking video.

  • Scripts are accepted when Sume estimates the video at 4 to 60 seconds; split longer scripts into several jobs.
  • aspect_ratio is 1:1, 3:4, 9:16 (default), 4:3 or 16:9, and 720p is the documented resolution.
  • In current code the avatar speaks English only. For other languages, generate speech with TTS 1.0 (POST /v1/tts-1.0/generate, with language set) and lip-sync it with Fabric.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume