Live AI avatar API: a Tavus conversation vs a Sume job

A live avatar API creates a room you join. Sume's Avatar API creates a job you poll. Field-by-field map of Tavus create conversation and Sume talking-video.

5 min readSume
All posts

A live AI avatar API creates a conversation: you call an endpoint, get a URL, and a person joins a room where the avatar talks back. Tavus's create-conversation endpoint works this way and returns a conversation_url (read 2026-10-03). Sume's Avatar API has no room: POST /v1/avatar-1.0/talking-video creates a job, returns polling URLs, and ends with a finished MP4 of 4-60 seconds.

If you are choosing between them, the quickest test is whether anything in your product needs to be said after the viewer arrives. If it does, you want a session API. If it does not, a job API is simpler to run.

What does a conversation request look like?

From Tavus's reference: you pass pal_id or face_id, and optional fields such as conversational_context, custom_greeting (spoken verbatim), callback_url, document_ids, meeting_url for Google Meet, Zoom or Teams, require_auth and policy: "eu". A nested properties object sets max_call_duration (default 3600 seconds), participant_absent_timeout (default 300), enable_recording, languages (up to 42) and recording_storage (S3, GCS or Azure Blob). The response carries conversation_id, conversation_url, status of active or ended, and, with require_auth, a short-lived meeting_token.

What does a Sume avatar request look like?

You pass an avatar_handle and exactly one of script or video_inputs. Optional fields are quality (standard, plus, max), aspect_ratio (default 9:16), scene, product_image, captions, and the communication fields mode and webhook_url. The Idempotency-Key header makes a retried submit safe. The response carries status_url, result_url, events_url and cancel_url, and result_ready tells you when GET result_url will work rather than return 409 job_not_completed (Generate avatar video, Jobs and results).

curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: welcome-clip-001" \
  -d '{
    "avatar_handle": "sume_clawra",
    "script": "Welcome aboard. Here is how to set up your workspace in three steps.",
    "quality": "standard",
    "aspect_ratio": "16:9",
    "mode": "async"
  }'

How do the two APIs map field by field?

The concepts line up loosely, and the differences are where integrations break. A conversation has a lifetime; a job has a terminal state. A greeting is set by a field; in Sume the opening is just your first scene.

Tavus create conversation vs Sume talking-video, read 2026-10-03
NeedTavus conversationSume talking-video
Who the avatar ispal_id or face_idavatar_handle
What it saysGenerated live; custom_greeting for the openerYour script or video_inputs
How you learn it is donestatus, callback_urlstatus_url, terminal, webhook_url
Lengthmax_call_duration, default 3600 s4-60 s estimated
Dry run or previewtest_mode creates a conversation with no PAL joiningavatar-video-previews then generate-video
Captionsenable_closed_captionsInline captions, burned into the MP4
What you keepOptional recording to your bucketA public media.sume.com video URL

What happens when something goes wrong?

A conversation ends on its own clocks. Tavus's reference sets participant_absent_timeout to 300 seconds by default, so a room nobody joins shuts down, and participant_left_timeout shuts it down after the last person leaves (read 2026-10-03). Your handling is mostly about the session: who rejoins, and what the person saw when it dropped.

A job fails once and says why. Sume's events_url returns sanitized lifecycle events, a job.failed webhook carries a public error, and cancel_url works only while cancelable is true, which is before generation starts. Replaying a submit with the same Idempotency-Key returns the original with idempotency_hit rather than creating a second paid job, which is the property that makes retries safe in a queue worker (Jobs and results).

Which fields have no counterpart?

Several Tavus features exist because a conversation is live, and Sume has nothing to map them to. document_ids and document_retrieval_strategy let a PAL read your documents while it talks; participant_tags keeps memory across calls; meeting_url sends a PAL into Google Meet, Zoom or Teams; require_auth and meeting_token gate who joins (Tavus reference, read 2026-10-03). In a job API, the equivalent of grounding is writing the right facts into the script before you submit.

The reverse is also true. Sume's video_inputs scenes, silence beats, product_image and scene references, and inline captions have no meaning in a conversation, because they describe a finished edit. If your brief reads like an edit decision list, you are describing a clip.

Which should my backend use?

Build on a conversation API when the call itself is the product. Build on a job API when the video is an asset you will attach to an email, a page or an app screen. With Sume, mode: "sync" waits at most 30 seconds, so for video submit async and poll or take the webhook; the status and result calls are the same for every Sume model.

  • Session API: a lifetime, a join URL, per-minute usage, someone must be present.
  • Job API: a terminal state, an artifact URL, per-clip usage, nobody needs to be present.
  • Mixed: render the fixed answers as clips and let a live agent handle the rest.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume