bitHuman local .imx or cloud avatar vs a hosted Sume avatar job
bitHuman renders avatars locally from .imx files or in its cloud, which changes your CPU bill. A hosted Sume avatar job moves the render off your agent.
bitHuman's LiveKit plugin can render a realtime avatar either locally from an .imx model file or in the cloud from an avatar ID, and the docs warn that agents use more CPU with .imx files. A hosted Sume avatar job sidesteps the question: the render runs on Sume workers and your process only submits and polls.
What the plugin offers
The LiveKit page says bitHuman avatars run "either locally or in the cloud". The plugin accepts three setups: an .imx model file from bitHuman's console, a direct image (a PIL object, a file path or a URL), or an existing bitHuman avatar ID. You install it with uv add "livekit-agents[bithuman]~=1.8" and set BITHUMAN_API_SECRET, plus an optional BITHUMAN_MODEL_PATH.
AvatarSession takes a model of "expression" or "essence". Essence is the default and gives predefined actions and expressions.
Where the work runs
Local rendering keeps media on your machines but turns your agent process into a render host, so you size CPU for it. Cloud rendering moves that load away at the price of a vendor dependency.
With Sume, you never hold a model file. Create an avatar once, then send avatar_handle and a script to POST /v1/avatar-1.0/talking-video. Sume also takes quality (standard, plus, max), aspect_ratio (1:1, 3:4, 9:16, 4:3, 16:9; default 9:16) and resolution (720p).
| Item | bitHuman local `.imx` | bitHuman cloud ID | Sume avatar job |
|---|---|---|---|
| Render host | Your agent's CPU | bitHuman's cloud | Sume workers |
| Setup | Download .imx from the console | Pick an avatar ID | Create an avatar, keep the handle |
| Output | Live avatar in a room | Live avatar in a room | MP4 job result |
| Secret | BITHUMAN_API_SECRET | BITHUMAN_API_SECRET | SUME_API_KEY |
| Duration | Session length | Session length | 4 to 60 s per video |
A practical way to choose
Ask who pays for idle time. A live avatar occupies compute for as long as the room is open. A Sume job occupies compute only while it renders, then the result is a file that costs nothing to replay.
If your script is fixed, render it once. If viewers drive the conversation, use a live plugin and budget the CPU. Our avatar versus lip-sync versus motion-control guide shows which Sume route fits each input.
Sources
Related posts
More in Comparisons
- Can an AI avatar see me on a video call? Griffin perception vs Sume
Tavus says Griffin perceives visual context: 3.73 of 5 on VideoFDB perception vs 4.20 for humans. Sume avatar clips cannot see you, but a clip can be inspected.
- Can gpt-live-transcribe take an uploaded file? Endpoints vs Sume STT
OpenAI lists gpt-live-transcribe at $0.017/min and only for the realtime transcription endpoint. For a file, Sume STT is $0.01/min, up to 10 minutes per job.
- Cartesia Ink at $0.39 an hour vs ElevenLabs Scribe v2 at $0.22
Cartesia lists Ink STT at $0.39 an hour on the Scale plan; ElevenLabs lists Scribe v2 at $0.22. Ink costs 1.77x as much; Sume's STT is $0.01 a minute.
- Clean vs verbatim transcripts: MAI-Transcribe-2 style vs Sume captions
MAI-Transcribe-2 batch has transcribeStyle clean or verbatim. Sume STT has no style flag; for polished captions supply script_text and keep your wording.
Written by Sume