Does Sume have a real-time avatar API? No, here is what it has instead

Sume has no live avatar session. It has async avatar jobs: create an avatar, render a talking video, or lip-sync a still to audio. Routes, limits and prices.

4 min readSume
All posts

No. Sume does not offer a real-time avatar API. There is no live session, no microphone input, and no stream of video frames that answers a viewer as they talk. Sume avatars are async jobs: you submit a request, get a job id back, and read the finished video later. There is no live video session. What Sume offers is a set of routes that make finished avatar videos, each as a job.

If you came here after reading about live conversational video from other companies, this page lists what you can build on Sume today and where the limits are.

What Sume offers

Avatar routes on Sume (Sume docs and provider-pricing, read 2026-10-07)
RouteWhat it doesLimitPrice
POST /v1/avatar-1.0/generateCreates a reusable avatar from a prompt, profile or photoOne job per avatar$0.95 flat per avatar
POST /v1/avatar-1.0/talking-videoScript or scenes to a talking video4 to 60 seconds, 720p$0.184 to $0.55 per second by tier
POST /v1/veed/fabric-1.0Still plus audio to a talking clipDefault lip-sync model$0.10 per second at 480p, $0.1875 at 720p
POST /v1/minimax/h3-max/lip-syncStill plus audio to a talking clipAudio 5 to 14.8 seconds$0.0625 to $0.20 per second by resolution

How a job runs

You submit, and Sume returns 202 with a job id and poll URLs. The job is queued, then processing, then completed, failed or canceled. You poll the status URL, or you send mode: "webhook" with a public HTTPS webhook_url and Sume posts a terminal event. The docs say that video and avatar-video jobs usually take longer than the 30-second wait that sync mode allows, so plan for waiting, not for a blocking call.

What you can do instead of live

Most uses of a live avatar have an async version. A greeting can be a pre-rendered clip for each visitor segment. A support answer can be a short clip written for the top questions. A sales follow-up can be a recap clip sent after the call. A launch can be a spokesperson video with captions, made at the quality tier you pick. For each, the clip is a file that you can review, caption, host, and reuse.

What does not move over is the open conversation. If the viewer's own words must change what the avatar says next, you need a live product, and you can still use Sume for the prepared clips that surround it.

A quick test

Ask whether the content of the next sentence depends on something that has not happened yet. If yes, that sentence cannot be pre-rendered. If no, write it, review it, and render it. Most business video falls in the second group, and for that group the per-second price, the review step and the reuse make async the better deal.

  • Create one avatar for $0.95 and reuse its handle.
  • Render clips from 4 to 60 seconds, by script or by scene plan.
  • Pick standard, plus (the default) or max for each clip.
  • Use a preview to approve the first frame before the full render.
  • Wait with polling or a webhook; never block a page on a render.

A migration note for live prototypes

If you already have a live prototype, you do not need to throw it away. Keep the live layer for the conversation, and send the prepared moments through Sume: the opening clip, the product walkthrough, the recap. The two layers share a script and a persona, but they have separate timing. The live layer answers in a second. The Sume layer answers in the time of a job, which is minutes for a video, and that gap is the thing to design around.

Be transparent with viewers about which is which. A clip made by Sume is a rendered video of a synthetic presenter, and a label that says so protects the trust you built.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume