Beyond Presence 50 credits a minute vs Sume lip sync per second
Beyond Presence bills live speech-to-video at 50 credits a minute. Sume bills a rendered lip-sync clip per second. Here is how to compare the two units.

Beyond Presence prices its speech-to-video (S2V) API at 50 credits per minute and its managed agents API at 100 credits per minute. Sume prices a lip-sync clip per second of finished video, so a 10-second 768p clip costs $1.00. The two units do not convert, because one meters a live session and the other meters a file you keep.
Beyond Presence's page does not state a dollar value for a credit in the part I read, so this post does not turn its credits into dollars. Use it to decide which unit fits your product, then price your own volume.
What Beyond Presence lists
The page describes two products. The Speech-to-Video API turns an audio stream into a live avatar. The Managed Agents API bundles voice, face and vision into one agent. It lists about 100 ms latency for the foundational speech-to-video model, sub-1.2 s for managed agents, 1080p at 35 FPS, and nine stock avatars.
Those numbers matter if a person is waiting on a reply. Every minute the stream is open is metered, whether or not anyone speaks.
What Sume does instead
Sume does not stream. You send a still (or a ready avatar handle) plus a Sume-hosted audio file of 5 to 14.8 seconds to POST /v1/minimax/h3-max/lip-sync. You get a job, poll it, and download an MP4 you can reuse for every viewer.
The price is ceil(duration_seconds) times the per-second list rate for the resolution, times the 1.25 house margin. The API reserves that amount at admit, captures it on completion and refunds it on failure.
Side by side
The table puts the unit, the speed and the output of each next to each other.
| Question | Beyond Presence (live) | Sume (rendered) |
|---|---|---|
| Billing unit | Credits per minute: 50 for S2V, 100 for managed agents | Per second of output, ceil(duration_seconds) |
| Speed | About 100 ms (S2V model), sub-1.2 s (managed agents) | Async job: queued, processing, completed |
| Resolution | 1080p at 35 FPS | 480p, 768p (default), 1080p; no 2K |
| Input | Audio stream | Still or avatar handle plus Sume-hosted audio, 5 to 14.8 s |
| Example price | Credits; dollar value not on the page I read | 10 s at 768p: 10 x $0.08 x 1.25 = $1.00 |
| Reuse | A session ends when it ends | One MP4, played any number of times |
When each one wins
Pick a live avatar when the viewer asks a question you cannot predict and expects an answer in about a second. Pick a rendered clip when the script is known in advance and many people will watch the same file. Welcome messages, product explainers and weekly updates all fall in that second group, and you pay once instead of per viewer-minute.
If you need both, render the predictable parts as clips and keep the live agent for the open questions. See live minutes versus one render for the break-even arithmetic, and the per-second lip-sync price list for 480p to 1080p.
Sources
Related posts
More in Comparisons
- bitHuman local .imx or cloud avatar vs a hosted Sume avatar job
bitHuman renders avatars locally from .imx files or in its cloud, which changes your CPU bill. A hosted Sume avatar job moves the render off your agent.
- Can an AI avatar see me on a video call? Griffin perception vs Sume
Tavus says Griffin perceives visual context: 3.73 of 5 on VideoFDB perception vs 4.20 for humans. Sume avatar clips cannot see you, but a clip can be inspected.
- Can gpt-live-transcribe take an uploaded file? Endpoints vs Sume STT
OpenAI lists gpt-live-transcribe at $0.017/min and only for the realtime transcription endpoint. For a file, Sume STT is $0.01/min, up to 10 minutes per job.
- Cartesia Ink at $0.39 an hour vs ElevenLabs Scribe v2 at $0.22
Cartesia lists Ink STT at $0.39 an hour on the Scale plan; ElevenLabs lists Scribe v2 at $0.22. Ink costs 1.77x as much; Sume's STT is $0.01 a minute.
Written by Sume