Kling @Element character binding vs Sume: what to use instead

Kling v3 Pro on fal binds characters with elements. Sume's kling-3 takes frames only, so use a reference-capable model or one fixed first frame.

5 min readSume
All posts

Kling v3 Pro on fal has an elements input for custom characters and objects, referenced in the prompt as @Element1, @Element2. Sume's kling-3 has no equivalent: it takes a prompt with an optional start and end frame and no reference images; its catalog row says reference_images: false. For a recurring character on Sume, pick a model that accepts reference images, or lock the character in a first-frame still.

So the translation is not a field rename. It is a choice of model and workflow.

What does the fal page document for elements?

The fal API page describes elements as custom character or object support via image sets, a frontal image plus reference images from up to three angles, or via video. Elements are referenced in the prompt as @Element1, @Element2 and so on, and the page says they support voice binding. The page also lists generate_audio (default true), described as supporting Chinese and English voice output, with other languages translated to English.

That is everything I can confirm from the page I read. Whether elements keep a face stable across separate generations is not stated there.

Which Sume video models take references?

Read the live values from GET /v1/video-router/models rather than a post: the Seedance rows share common defaults that I did not list here, and capabilities change. On POST /v1/videos, the equivalent fields are supported_frame_images and supported_input_references in the video model list.

Video Router capability flags for reference and frame inputs (Sume catalog, read 2026-10-02)
Model idimage-to-videoend framereference images
kling-3YesYesNo
grok-imagine-video-1.5YesNoNo
wan-3.0YesYesYes
minimax-h3YesYesYes
minimax-h3-maxYesYesYes
gemini-omni-flash-1.1YesYesYes
seedance-2 family and seedance-2.5See the live model listSee the live model listSee the live model list

How do I carry a character without elements?

Two approaches work with what the docs describe. First, pick a reference-capable model and send the character image as an input_references entry. Only models whose supported_input_references lists a type accept it; a request that sends both frame_images and input_references is treated as image-to-video, because frame_images wins. Second, keep using kling-3 but start every shot from a still of the same character, made once and reused as each shot's first_frame.

The second approach pins appearance at the first frame only. Later frames can drift, which is why a frame check on each finished clip belongs in the loop; the post on consistent characters across shots goes into prompting for it.

{
  "model": "wan-3.0",
  "prompt": "The courier walks into the lobby and checks her phone",
  "duration": 8,
  "resolution": "720p",
  "aspect_ratio": "9:16",
  "input_references": [
    { "type": "image_url",
      "image_url": { "url": "https://media.sume.com/artifacts/artf_demo/courier-front.png" } },
    { "type": "image_url",
      "image_url": { "url": "https://media.sume.com/artifacts/artf_demo/courier-side.png" } }
  ]
}

What about voices tied to a character?

fal's page mentions voice binding on elements. Sume does not offer that in kling-3. Its agent guidance says video models do not lip-sync, so a character who speaks is made as speech first and then a talking-clip job from the approved still, and the audio becomes the spine of a Timeline 1.0 render. Keep the character's look in the still and the voice in the speech step; the two are joined at assembly, not inside one generation.

This is a trade, not a gap to hide. Bound elements are convenient inside one vendor's model. Sume's route is more steps, with the benefit that each step can be redone alone.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume