What is an AI video agent? One that talks, or one that makes

An AI video agent can mean a live on-camera persona like a Tavus PAL, or an agent that makes videos for you. Which one Sume is, and how to call it from code.

5 min readSume
All posts

"AI video agent" means two different things in search results. One is an agent that appears on camera and talks with you, such as a Tavus PAL or an Anam persona. The other is an agent that makes video for you: it plans a clip, calls generation tools, and returns a file. Sume is the second kind. It has no on-camera persona that talks back live.

Knowing which one you want saves a week of evaluating the wrong product.

What is the on-camera kind?

Tavus's documentation says it now uses PAL for what you build, Face for how it looks and Voice for how it sounds, with example PALs for sales, interviewing, customer support and medical intake (read 2026-10-03). Anam defines a persona as a face, a voice, an LLM and a system prompt, running as speech-to-text, LLM, text-to-speech and face generation (read 2026-10-03). Both put an agent on screen in a live conversation.

This is the kind of video agent that Tavus's Griffin extends into a single full-duplex model, currently in a research preview for trusted testers only (Tavus Griffin post).

What is the video-making kind, and what does Sume ship?

Sume's Agent runs the same runtime as the Agents chat in the app: sandbox, tools, media generation. You reach it from code with Agent Completions, POST /v1/agent/completions, which takes an instruction and returns an async agent.run receipt that you poll (Agent Completions). A Format stores a repeatable workflow, and a schedule runs a saved automation on a trigger.

The agent's output is media: images, video, music, narration and avatar videos. On hosted MCP, the avatar tools are avatars_create, avatar-videos_create, avatar-video-previews_create and the matching reads (MCP tools and gates). Paid tools need explicit write access and an idempotency key.

Two kinds of AI video agent, read 2026-10-03
QuestionTalks on camera (Tavus PAL, Anam persona)Makes video (Sume Agent)
What the viewer seesA live face in a callA finished clip
InteractionReal time, both waysYou send a task and read a receipt
How you call itA conversation API with a join URLAgent Completions, hosted MCP, or the job API
Cost shapeConversation minutesPer generation job
Review before outputNonePreviews, and confirmation before paid actions

What does a call to the making agent look like?

Send an instruction (or messages) and a required generation_spend_cap_usd, the most the run may spend on generation. The call returns 202 with a receipt that has a status_url and a cancel_url; you poll the run, and a replayed Idempotency-Key returns the original receipt. Each completion runs in a fresh thread (Agent Completions).

curl -X POST https://api.sume.com/v1/agent/completions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: video-brief-001" \
  -d '{
    "instruction": "Make a 20 second avatar clip that welcomes new customers and names the three setup steps.",
    "generation_spend_cap_usd": 5
  }'

How do I decide in five minutes?

Ask what the viewer does. If they talk to the agent, you want the on-camera kind, and you should compare vendors on latency, concurrency and per-minute pricing. If they watch something the agent made, you want the video-making kind, and you should compare on control: previews, review steps, spend caps and how results come back.

Search results blur the two because both are sold as video agents. A quick filter is the pricing unit. Per-minute of conversation means live; per-job or per-generation means a maker. Tavus's plans are per conversation minute (Tavus pricing, read 2026-10-03), while Sume's avatar renders are billed per job.

One more check is who is accountable for the output. An on-camera agent says things live, so your guardrails sit in its prompt and tools. A making agent produces a file you can open before anyone else sees it, so your guardrail is a review step and a spend cap. Sume's Agent Completions require generation_spend_cap_usd, and the same docs treat paid media creation as something to confirm, not assume.

Can the two be combined?

Yes, and this is a sensible split. Let the making agent produce the assets a live agent will point at: product explainers, intro clips, captioned answers to the questions asked every day. A live agent then covers the open questions. Sume's documentation says agents should read catalog, jobs and usage first and ask for confirmation before write or paid actions (Safe automation).

  • Need a face that answers questions: a live avatar vendor.
  • Need clips produced from a brief: a Sume Agent Completion or hosted MCP.
  • Need both: render the fixed answers first, run the live agent for the rest.
  • Always cap spend: an Agent Completion requires a generation spend cap.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume