LemonSlice API: image plus streaming audio vs Sume lip sync

LemonSlice drives a live avatar from an image and streaming audio. For a recorded line, Sume's lip sync takes a still plus an audio_url and returns a file.

4 min readSume
All posts

LemonSlice's two inputs are an image and audio, like a lip sync API, but the audio is streamed and the avatar answers in real time. Sume's still-plus-audio route is the recorded version: you send an image_url and a finished audio_url, and you get a job with a video file.

Which one you want depends on whether a person is talking back. This post sets the two side by side using LemonSlice's docs and Sume's models page.

What does LemonSlice take as input?

Its docs list an image that defines the character, which can be created from a single photo, and audio streamed from text to speech services such as ElevenLabs or Cartesia. An optional Action Engine for gestures is marked Enterprise only. The pipeline runs speech to text, voice activity detection, an LLM, text to speech and video generation over WebRTC.

The page says it supports 1000+ simultaneous calls and multi-hour conversations. It does not give pricing or resolution limits, so this post states none.

What does Sume take for a recorded line?

VEED Fabric 1.0 is public as veed/fabric-1.0 at POST /v1/veed/fabric-1.0. Send audio_url, a measured duration_seconds, and exactly one visual source: an image_url or an avatar_handle. They are mutually exclusive. Sume's guidance is to prefer the generated, inspected still as the image_url, and to use the handle only when the user named that avatar.

MiniMax H3 Max Lip Sync, at POST /v1/minimax/h3-max/lip-sync, takes the same body for audio between 5 and 14.8 seconds. Video models do not lip-sync to generated speech laid on afterward, which is why the talking face goes through a still-plus-audio model.

LemonSlice from its docs page; Sume from the Models page; both read 2026-10-03.
QuestionLemonSliceSume Fabric / H3 Max lip sync
Visual inputAn imageimage_url or avatar_handle, not both
Audio inputStreaming TTS audioA finished audio_url plus duration_seconds
InteractionListens and answers in real timeNone: renders the audio you give it
OutputA live sessionA job, then a video file
Scale noted1000+ simultaneous callsPlan queue limits, set by workspace

Where does the audio come from?

On LemonSlice, a live pipeline produces it as the conversation goes. On Sume, you make it first, usually with a text to speech job, then pass its URL. That keeps the voice reviewable before you pay for video, and a clean take can be reused.

Do not stretch or re-cut audio just to fit a model. H3 Max's window is 5-14.8 seconds and Sume refuses a duration_seconds outside it rather than clamping; a clip outside the window goes to Fabric.

Which should you pick?

Pick LemonSlice-style real-time when the avatar must respond to a live person. Pick a rendered still-plus-audio job when the line is known, you want it checked, and you want a file you can publish. Both start from one picture and one voice; only one of them has a listener.

Submit rendered jobs with an Idempotency-Key, keep the job.id, and poll status. A wait you spend on a long clip should not trigger a second paid submit.

What goes wrong with still-plus-audio jobs?

Two checks catch most failures. The audio must live on a Sume-hosted URL and stay within the size limit, or the submit is refused with unsupported_audio_source or audio_too_large; and the duration_seconds you send must match the real length of the audio, because it is the basis for the price reserved at submit.

The reservation is captured when the job completes and refunded if it fails, so a wrong duration shows up as a surprise on the bill, not as an error. Measure the audio, then send the number.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume