What is an AI video agent? One that talks, or one that makes
An AI video agent can mean a live on-camera persona like a Tavus PAL, or an agent that makes videos for you. Which one Sume is, and how to call it from code.

"AI video agent" means two different things in search results. One is an agent that appears on camera and talks with you, such as a Tavus PAL or an Anam persona. The other is an agent that makes video for you: it plans a clip, calls generation tools, and returns a file. Sume is the second kind. It has no on-camera persona that talks back live.
Knowing which one you want saves a week of evaluating the wrong product.
What is the on-camera kind?
Tavus's documentation says it now uses PAL for what you build, Face for how it looks and Voice for how it sounds, with example PALs for sales, interviewing, customer support and medical intake (read 2026-10-03). Anam defines a persona as a face, a voice, an LLM and a system prompt, running as speech-to-text, LLM, text-to-speech and face generation (read 2026-10-03). Both put an agent on screen in a live conversation.
This is the kind of video agent that Tavus's Griffin extends into a single full-duplex model, currently in a research preview for trusted testers only (Tavus Griffin post).
What is the video-making kind, and what does Sume ship?
Sume's Agent runs the same runtime as the Agents chat in the app: sandbox, tools, media generation. You reach it from code with Agent Completions, POST /v1/agent/completions, which takes an instruction and returns an async agent.run receipt that you poll (Agent Completions). A Format stores a repeatable workflow, and a schedule runs a saved automation on a trigger.
The agent's output is media: images, video, music, narration and avatar videos. On hosted MCP, the avatar tools are avatars_create, avatar-videos_create, avatar-video-previews_create and the matching reads (MCP tools and gates). Paid tools need explicit write access and an idempotency key.
| Question | Talks on camera (Tavus PAL, Anam persona) | Makes video (Sume Agent) |
|---|---|---|
| What the viewer sees | A live face in a call | A finished clip |
| Interaction | Real time, both ways | You send a task and read a receipt |
| How you call it | A conversation API with a join URL | Agent Completions, hosted MCP, or the job API |
| Cost shape | Conversation minutes | Per generation job |
| Review before output | None | Previews, and confirmation before paid actions |
What does a call to the making agent look like?
Send an instruction (or messages) and a required generation_spend_cap_usd, the most the run may spend on generation. The call returns 202 with a receipt that has a status_url and a cancel_url; you poll the run, and a replayed Idempotency-Key returns the original receipt. Each completion runs in a fresh thread (Agent Completions).
curl -X POST https://api.sume.com/v1/agent/completions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: video-brief-001" \
-d '{
"instruction": "Make a 20 second avatar clip that welcomes new customers and names the three setup steps.",
"generation_spend_cap_usd": 5
}'How do I decide in five minutes?
Ask what the viewer does. If they talk to the agent, you want the on-camera kind, and you should compare vendors on latency, concurrency and per-minute pricing. If they watch something the agent made, you want the video-making kind, and you should compare on control: previews, review steps, spend caps and how results come back.
Search results blur the two because both are sold as video agents. A quick filter is the pricing unit. Per-minute of conversation means live; per-job or per-generation means a maker. Tavus's plans are per conversation minute (Tavus pricing, read 2026-10-03), while Sume's avatar renders are billed per job.
One more check is who is accountable for the output. An on-camera agent says things live, so your guardrails sit in its prompt and tools. A making agent produces a file you can open before anyone else sees it, so your guardrail is a review step and a spend cap. Sume's Agent Completions require generation_spend_cap_usd, and the same docs treat paid media creation as something to confirm, not assume.
Can the two be combined?
Yes, and this is a sensible split. Let the making agent produce the assets a live agent will point at: product explainers, intro clips, captioned answers to the questions asked every day. A live agent then covers the open questions. Sume's documentation says agents should read catalog, jobs and usage first and ask for confirmation before write or paid actions (Safe automation).
- Need a face that answers questions: a live avatar vendor.
- Need clips produced from a brief: a Sume Agent Completion or hosted MCP.
- Need both: render the fixed answers first, run the live agent for the rest.
- Always cap spend: an Agent Completion requires a generation spend cap.
Sources
Related posts
More in Agents
- Run the Sume video agent from your backend with Agent Completions
POST /v1/agent/completions runs the same agent as the Sume Agents chat, with tools and media generation, and returns an async run receipt you poll or webhook.
- Safe automation for AI agents that call paid APIs
Keep agents read-only by default, keep secrets out of logs, and on hosted MCP send an idempotency_key, preview with dry_run, and cap with max_spend_usd.
- Scheduled AI video agent runs: cron, API triggers, and receipts
A Sume schedule is a saved Agents automation that runs on a cron cadence and returns a run receipt. Author it in the dashboard; start and monitor runs by API.
- What is a video agent? How Sume defines and runs one
In Sume's docs, a video agent is a sandbox Agent that composes generation tools into a post-ready video. Brief it in chat, or call it over HTTP.
Written by Sume