Is AI avatar video real time? How long a Sume job takes

A Sume avatar video is a job, not a live stream: it queues, renders, and you poll or take a webhook. What the sync wait caps at, and a Python polling loop.

5 min readSume
All posts

No. A Sume avatar video is not real time: you submit a script, Sume queues a job, renders the clip, and you read the finished MP4 afterwards. The Sume docs give no render time for avatar video, so treat it as an asynchronous wait of unknown length and build for polling or a webhook. The only fixed number in the docs is the HTTP wait: sync blocks the submit call for at most 30 seconds.

Live avatars are a different category. Tavus's platform page reports 134 ms for its Phoenix-4.5 rendering model and its Griffin post reports 0.43 seconds of video latency on H100s (both read 2026-10-03). Those figures describe a stream where frames arrive as they are drawn. A job returns when the whole clip is done.

What happens between submit and result?

A job moves through queued, processing, then a terminal completed, failed or canceled. A queued job is a normal accepted state: workspace concurrency limits apply when workers move jobs into processing, not when the API accepts a valid request (Jobs and results).

The submit response carries status_url, result_url, events_url and, until generation starts, cancel_url. The status payload adds terminal, result_ready, next_poll_after_seconds and recommended_poll_interval_seconds, so your client does not need to guess a polling cadence. GET result_url returns 409 job_not_completed until result_ready is true.

Ways to learn an avatar job finished, from the Sume docs, read 2026-10-03
ModeWhat the submit call doesWait on the HTTP callUse it for
async (default)Returns 202 with polling URLsNonePoll status_url
webhookReturns the job; Sume posts job.completed laterNoneServer-to-server delivery
syncWaits for a terminal stateAt most 30 sShort jobs; still poll if not terminal
subscribeAlias of syncAt most 30 sSame as sync

Can I make it faster?

Choose the quality tier. standard is described as the fastest Sume execution path, plus is the default, and max is the highest quality with slower turnaround (Generate avatar video). The docs give no timing per tier, so measure on your own account.

The bigger lever is when you submit. Render ahead of the moment a viewer needs the clip, and store the result URL. A welcome clip rendered when the account is created is ready before the user opens the app.

How do I wait for it in Python?

Submit with async, poll the status URL on the interval Sume recommends, then fetch the result once the job is terminal and completed. Never resubmit a job that has not finished; the 30-second sync ceiling bounds the HTTP wait, not the job.

import os, time, requests

BASE = "https://api.sume.com/v1/avatar-1.0/talking-video"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
r = requests.post(
    BASE,
    headers={**H, "Idempotency-Key": "wait-demo-001"},
    json={"avatar_handle": "sume_clawra", "quality": "standard",
          "script": "Your report is ready. Open the dashboard to see this week's numbers.",
          "mode": "async"},
    timeout=30,
)
r.raise_for_status()
job = r.json()["data"]
while True:
    s = requests.get(job["status_url"], headers=H, timeout=30).json()["data"]
    if s["terminal"]:
        break
    time.sleep(s.get("recommended_poll_interval_seconds") or 5)
if s["sume_status"] != "completed":
    raise SystemExit(f"job ended as {s['sume_status']}")
print(requests.get(job["result_url"], headers=H, timeout=30).json())

What if the job fails or I need to stop it?

Read sume_status: failed and canceled are terminal, and events_url gives sanitized lifecycle events with the reason. cancel_url works only while cancelable is true, which is before generation has started; after that it is null. If the job completed but your webhook receiver was down, the API has a POST /v1/jobs/{id}/webhook/redeliver route for resending the terminal event, so a missed delivery does not mean a lost clip.

On any retry, reuse the same Idempotency-Key. A replayed submit returns the original job with idempotency_hit set, instead of starting a second paid render.

Where does real time still matter for clips?

Mostly at the edges of your product. If a user presses a button and expects a talking presenter at once, a job is the wrong tool; render ahead, or render a small set of likely clips and pick one. If the clip is for a campaign, a catalog or a help page, wait time disappears because rendering happens before anyone watches.

This is why a common architecture is a hybrid: a live agent for the conversation, and rendered clips for anything you can predict. The clips are then served like any other video file and cost the live session nothing in latency.

What if I really need a live avatar?

Then a job is the wrong tool. Use a session product for the conversation and keep Sume for the assets around it: intro clips, explainer answers you give every time, captioned recordings. The two work well together, because a clip that is ready before the call costs the call nothing in latency.

  • Real time needed: a live avatar vendor, per-minute pricing.
  • Wait acceptable, quality matters: a rendered clip, per-render pricing.
  • Never block a user-facing request on a render; hand the user a placeholder and notify on completion.
  • Use an Idempotency-Key on every submit so retries do not create duplicates.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume