Kling @Element character binding vs Sume: what to use instead
Kling v3 Pro on fal binds characters with elements. Sume's kling-3 takes frames only, so use a reference-capable model or one fixed first frame.

Kling v3 Pro on fal has an elements input for custom characters and objects, referenced in the prompt as @Element1, @Element2. Sume's kling-3 has no equivalent: it takes a prompt with an optional start and end frame and no reference images; its catalog row says reference_images: false. For a recurring character on Sume, pick a model that accepts reference images, or lock the character in a first-frame still.
So the translation is not a field rename. It is a choice of model and workflow.
What does the fal page document for elements?
The fal API page describes elements as custom character or object support via image sets, a frontal image plus reference images from up to three angles, or via video. Elements are referenced in the prompt as @Element1, @Element2 and so on, and the page says they support voice binding. The page also lists generate_audio (default true), described as supporting Chinese and English voice output, with other languages translated to English.
That is everything I can confirm from the page I read. Whether elements keep a face stable across separate generations is not stated there.
Which Sume video models take references?
Read the live values from GET /v1/video-router/models rather than a post: the Seedance rows share common defaults that I did not list here, and capabilities change. On POST /v1/videos, the equivalent fields are supported_frame_images and supported_input_references in the video model list.
| Model id | image-to-video | end frame | reference images |
|---|---|---|---|
| kling-3 | Yes | Yes | No |
| grok-imagine-video-1.5 | Yes | No | No |
| wan-3.0 | Yes | Yes | Yes |
| minimax-h3 | Yes | Yes | Yes |
| minimax-h3-max | Yes | Yes | Yes |
| gemini-omni-flash-1.1 | Yes | Yes | Yes |
| seedance-2 family and seedance-2.5 | See the live model list | See the live model list | See the live model list |
How do I carry a character without elements?
Two approaches work with what the docs describe. First, pick a reference-capable model and send the character image as an input_references entry. Only models whose supported_input_references lists a type accept it; a request that sends both frame_images and input_references is treated as image-to-video, because frame_images wins. Second, keep using kling-3 but start every shot from a still of the same character, made once and reused as each shot's first_frame.
The second approach pins appearance at the first frame only. Later frames can drift, which is why a frame check on each finished clip belongs in the loop; the post on consistent characters across shots goes into prompting for it.
{
"model": "wan-3.0",
"prompt": "The courier walks into the lobby and checks her phone",
"duration": 8,
"resolution": "720p",
"aspect_ratio": "9:16",
"input_references": [
{ "type": "image_url",
"image_url": { "url": "https://media.sume.com/artifacts/artf_demo/courier-front.png" } },
{ "type": "image_url",
"image_url": { "url": "https://media.sume.com/artifacts/artf_demo/courier-side.png" } }
]
}What about voices tied to a character?
fal's page mentions voice binding on elements. Sume does not offer that in kling-3. Its agent guidance says video models do not lip-sync, so a character who speaks is made as speech first and then a talking-clip job from the approved still, and the audio becomes the spine of a Timeline 1.0 render. Keep the character's look in the still and the voice in the speech step; the two are joined at assembly, not inside one generation.
This is a trade, not a gap to hide. Bound elements are convenient inside one vendor's model. Sume's route is more steps, with the benefit that each step can be redone alone.
Sources
Related posts
More in Comparisons
- Kling v3 Pro multi_prompt shots on fal vs Sume's one clip per shot
fal's Kling v3 Pro has a multi_prompt list for multi-shot video. Sume's kling-3 sends one prompt and frames per clip, so you join shots with Timeline 1.0.
- Kling v3 Turbo Pro at $0.14 a second vs v3 Pro and Sume's kling-3
fal lists Kling v3 Turbo Pro at $0.14 a second and v3 Pro at $0.112 audio off, $0.168 on. What that means for 5 seconds, and the Kling id Sume lists.
- Leonardo.Ai API alternative for image generation: Sume Images
Leonardo's API submits to /generations and returns a generationId. Sume POST /v1/images works from a model catalog with idempotent retries. Compared.
- Live vs file transcription cost: gpt-live-transcribe and batch rates
Streaming transcription costs 2 to 4 times file transcription on vendor pages. When live is worth it, and when Sume's $0.01 per minute file job is enough.
Written by Sume