Kling Avatar v2 audio formats vs Sume Fabric audio_url

fal lists MP3, OGG, WAV, M4A and AAC audio for Kling Avatar v2. Sume Fabric takes an audio_url plus a still; here is how the two inputs compare.

4 min readSume
All posts

fal's Kling AI Avatar v2 accepts one audio file in MP3, OGG, WAV, M4A or AAC, plus one image. On Sume the closest routes are POST /v1/veed/fabric-1.0 and the MiniMax H3 Max lip sync: both take a still and an audio URL, and the Sume pages I read do not list audio file types.

Vendor facts are from fal's Kling AI Avatar v2 Pro page, read 2026-10-01. Sume facts are from the Models overview and Create new avatar.

Which file types does fal accept for Kling Avatar v2?

fal's input hints list audio as mp3, ogg, wav, m4a and aac, and images as jpg, jpeg, png, webp, gif, avif, heic and heif; the spec table on the same page lists JPG, JPEG, PNG, WebP, GIF and AVIF. The output is an MP4 with synchronized audio.

Kling AI Avatar v2 inputs on fal, read 2026-10-01.
InputTypes on fal
AudioMP3, OGG, WAV, M4A, AAC
Image (spec table)JPG, JPEG, PNG, WebP, GIF, AVIF
OutputMP4 with synchronized audio

What does Sume Fabric take?

The models overview lists VEED Fabric 1.0 at POST /v1/veed/fabric-1.0 as a talking still plus audio clip route, and the H3 Max lip sync as the same still and audio body with audio of 5 to 14.8 seconds. The Fabric body carries audio_url, a measured duration_seconds, and exactly one visual source: image_url or avatar_handle.

Those pages do not enumerate accepted audio containers. The OpenAPI document says audio_url must be a public HTTPS URL on the Sume media host (non-Sume hosts are rejected) and at most 10 MB, so upload or generate the audio on Sume first, unlike fal, which also accepts a URL you host. If your audio is in an unusual format, test it on a small job or read the live OpenAPI document before building on it.

What are the rules for the image?

For avatar creation, the docs say image_url must be a fetchable public HTTPS image URL, and localhost, private-network, non-HTTPS and non-image URLs are rejected before generation. Plan the same for Fabric stills: host the file publicly first.

Can I use generated speech instead of a recording?

Yes, as audio you produce first and then pass as the audio file. The docs are explicit that video models do not lip-sync to generated TTS or to a later voice-over, so a talking face goes through Fabric with a still and audio rather than a video clip with narration laid underneath. For the clip-length side of the same question see talking photo audio between 5 and 14 seconds.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume