LemonSlice API: image plus streaming audio vs Sume lip sync
LemonSlice drives a live avatar from an image and streaming audio. For a recorded line, Sume's lip sync takes a still plus an audio_url and returns a file.

LemonSlice's two inputs are an image and audio, like a lip sync API, but the audio is streamed and the avatar answers in real time. Sume's still-plus-audio route is the recorded version: you send an image_url and a finished audio_url, and you get a job with a video file.
Which one you want depends on whether a person is talking back. This post sets the two side by side using LemonSlice's docs and Sume's models page.
What does LemonSlice take as input?
Its docs list an image that defines the character, which can be created from a single photo, and audio streamed from text to speech services such as ElevenLabs or Cartesia. An optional Action Engine for gestures is marked Enterprise only. The pipeline runs speech to text, voice activity detection, an LLM, text to speech and video generation over WebRTC.
The page says it supports 1000+ simultaneous calls and multi-hour conversations. It does not give pricing or resolution limits, so this post states none.
What does Sume take for a recorded line?
VEED Fabric 1.0 is public as veed/fabric-1.0 at POST /v1/veed/fabric-1.0. Send audio_url, a measured duration_seconds, and exactly one visual source: an image_url or an avatar_handle. They are mutually exclusive. Sume's guidance is to prefer the generated, inspected still as the image_url, and to use the handle only when the user named that avatar.
MiniMax H3 Max Lip Sync, at POST /v1/minimax/h3-max/lip-sync, takes the same body for audio between 5 and 14.8 seconds. Video models do not lip-sync to generated speech laid on afterward, which is why the talking face goes through a still-plus-audio model.
| Question | LemonSlice | Sume Fabric / H3 Max lip sync |
|---|---|---|
| Visual input | An image | image_url or avatar_handle, not both |
| Audio input | Streaming TTS audio | A finished audio_url plus duration_seconds |
| Interaction | Listens and answers in real time | None: renders the audio you give it |
| Output | A live session | A job, then a video file |
| Scale noted | 1000+ simultaneous calls | Plan queue limits, set by workspace |
Where does the audio come from?
On LemonSlice, a live pipeline produces it as the conversation goes. On Sume, you make it first, usually with a text to speech job, then pass its URL. That keeps the voice reviewable before you pay for video, and a clean take can be reused.
Do not stretch or re-cut audio just to fit a model. H3 Max's window is 5-14.8 seconds and Sume refuses a duration_seconds outside it rather than clamping; a clip outside the window goes to Fabric.
Which should you pick?
Pick LemonSlice-style real-time when the avatar must respond to a live person. Pick a rendered still-plus-audio job when the line is known, you want it checked, and you want a file you can publish. Both start from one picture and one voice; only one of them has a listener.
Submit rendered jobs with an Idempotency-Key, keep the job.id, and poll status. A wait you spend on a long clip should not trigger a second paid submit.
What goes wrong with still-plus-audio jobs?
Two checks catch most failures. The audio must live on a Sume-hosted URL and stay within the size limit, or the submit is refused with unsupported_audio_source or audio_too_large; and the duration_seconds you send must match the real length of the audio, because it is the basis for the price reserved at submit.
The reservation is captured when the job completes and refunded if it fails, so a wrong duration shows up as a surprise on the bill, not as an error. Measure the audio, then send the number.
Sources
Related posts
More in Sume Avatar 1.0
- Lip sync audio under 5 seconds: H3 Max rejects it, use Fabric
Sume's MiniMax H3 Max lip sync only accepts 5 to 14.8 seconds of audio and returns invalid_request outside it. A 4.5-second line goes to Fabric instead.
- Lip sync looks fake? Check teeth, profile, timing and seams on Sume
sync. labs names four tells of fake lip sync. Pull PNG stills from a Sume avatar clip with video frames at the moments each tell shows up, then decide.
- Seasonal avatar host from a prompt, profile or photo for Christmas
Create a reusable Sume Avatar 1.0 handle from a prompt, structured profile props or a reference photo, then reuse it across holiday clips.
- How to make an AI UGC ad look less staged with Avatar 1.0
Less-staged AI UGC comes from the first frame: phone-style framing, a casual scene prompt, an approved preview, then the final render. The levers Sume exposes.
Written by Sume