LiveKit avatar providers with Node.js support vs a Sume Node job
LiveKit lists eight avatar providers with Node.js plugins and eight Python-only ones. For a clip rather than a live room, Sume works from Node with fetch.
LiveKit's avatar models page lists 16 providers, all with Python plugins, and eight with Node.js plugins: Anam, Beyond Presence, D-ID, LemonSlice, Protoface, Runway, Tavus and TruGen (read 2026-10-05). If your backend is Node and you want a prepared clip instead of a live participant, Sume's job API needs only fetch.
The support matrix
The other eight (Avatario, AvatarTalk, bitHuman, Keyframe, LiveAvatar, Simli, Spatius and Synthesia) are Python-only on that page. Check the page again before you commit, because the table changes as plugins ship.
| Provider group | Python | Node.js |
|---|---|---|
| Anam, Beyond Presence, D-ID, LemonSlice | Yes | Yes |
| Protoface, Runway, Tavus, TruGen | Yes | Yes |
| Avatario, AvatarTalk, bitHuman, Keyframe | Yes | No |
| LiveAvatar, Simli, Spatius, Synthesia | Yes | No |
A Node job instead of a plugin
A Sume lip-sync call is plain HTTPS. Submit with an Idempotency-Key, poll the status endpoint until a terminal state, then read the result. This needs Node 18 or later, run as an ES module for the top-level await, or wrap it in a function.
const API = "https://api.sume.com";
const headers = {
Authorization: `Bearer ${process.env.SUME_API_KEY}`,
"Content-Type": "application/json",
};
async function call(method, path, body, key) {
const h = key ? { ...headers, "Idempotency-Key": key } : headers;
const res = await fetch(API + path, {
method, headers: h, body: body ? JSON.stringify(body) : undefined,
});
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
return res.json();
}
const sub = await call("POST", "/v1/minimax/h3-max/lip-sync", {
image_url: process.env.IMAGE_URL,
audio_url: process.env.AUDIO_URL,
duration_seconds: 5.84,
resolution: "768p",
}, "node-lipsync-001");
let delay = 3000;
for (;;) {
const st = await call("GET", `/v1/jobs/${sub.job.id}/status`);
const state = (st.job ?? st).status;
if (["completed", "failed", "canceled"].includes(state)) break;
await new Promise((r) => setTimeout(r, delay));
delay = Math.min(delay * 2, 30000);
}
console.log(await call("GET", `/v1/jobs/${sub.job.id}/result`));Which to pick
Use a LiveKit plugin when the person on the other end must talk back in real time. Use a Sume job when the answer can be prepared: a greeting, a product explainer, a status update. The audio must be hosted on Sume media, and for H3 Max lip sync it must run 5 to 14.8 seconds. See measuring join and playback latency for judging the live side.
Sources
Related posts
More in Developers
- Load-test Sume job polling with k6: reads have their own budget
A k6 script that polls one finished Sume job from 5 virtual users, and what 4,800 reads a minute on Free means for any 429 you see.
- Log usage.cost for every Sume video job in Python and total a batch
A finished Sume video job returns usage.cost in USD. Log it with the job id and model, and sum it for a batch. Here is a short Python script that does it.
- Log usage.cost from the Sume images response to CSV in Python
POST /v1/images returns usage.cost as the billed USD amount. A Python logger that writes model, image count and cost per image to a CSV for budget reviews.
- Longest AI video Sume can render: 1800 s from 60 clips
One Timeline render caps at 1800 seconds. With 30-second clips that is 60 clips, with 200 slots allowed. The math and the render price follow.
Written by Sume