Video agent API: turn a brief into a finished, edited video

Yes: a video agent API turns a brief into a finished, edited video. What Sume's Agent Completions and Formats take, return, cost, and won't do.

6 min readSume
All posts

Yes. A video agent API takes a brief in plain words and returns a finished video: the agent plans the shots, calls generation models, and edits the pieces into one file, so you never write the timeline yourself. Sume documents two such calls: Agent Completions (POST /v1/agent/completions) for a one-off brief, and Formats (POST /v1/formats/{handle}/{slug}/runs) for a saved recipe you call with new inputs. HeyGen also documents a Video Agent endpoint.

The Sume facts come from the Agent Completions, Format API, and Sume basics docs; other vendors' facts come from their own developer docs. All were read on 2026-09-28. Below: what "finished, edited" means in those docs, how an agent API differs from a model or render API, a minimal request, and the limits.

What does "finished, edited" mean here?

Sume's basics page says the main partner path is a sandbox Agent or Format that composes the individual generation tools, which is how it ships what a single clip cannot: "multi-minute host footage, B-roll, voiceover, and timeline assembly into a post-ready video." The Format API run diagram lists the same work: host takes, B-roll, voiceover, captions, timeline assembly. What comes back:

  • One deliverable. A completed Format run carries primary_output_url, which the docs call the one thing to show, plus artifacts[], every file the run made.
  • Media in output. A completed Agent Completion puts the agent's closing text in output.text and generated media in output.videos, output.images, output.audio, and output.files.
  • Durable files. Media URLs are durable media.sume.com HTTPS URLs you can store.
  • All or nothing. A Format run is never a partial delivery: a run that could not finish comes back failed, and its primary_output_url is null.
  • Extras by request. The docs' own example instruction turns extras off by name ("No BGM, no captions"), so say in the brief whether you want voiceover, captions, and music.

How is an agent API different from a model or render API?

By how much you write. A model API makes one clip from a prompt. A render API edits clips you already have, along a timeline you write. An avatar API turns a script into a talking clip. An agent API takes the brief and does the planning, generation, and edit itself. Sume has an endpoint for each approach:

From Video Router, Timeline 1.0, Generate avatar video, Agent Completions, and Format API, read 2026-09-28.
ApproachWhat you sendWhat you get backSume endpoint
Raw model APIA prompt, plus options such as aspect_ratioOne clip; sume/auto makes 3–10 s clips at 16:9 or 9:16POST /v1/videos
Render (timeline) APIYour timeline: one audio spine plus 1–200 video slots with start times, all Sume-hostedOne MP4, 1080×1920 by defaultPOST /v1/timeline-1.0/render
Avatar APIA ready avatar and a scriptA talking video of 4–60 s at 720pPOST /v1/avatar-1.0/talking-video
Agent API, one-offA brief as instruction or messages, plus a spend capA run receipt, then finished media in outputPOST /v1/agent/completions
Agent API, saved recipeinstruction and input for a saved Formatprimary_output_url, artifacts[], and outputPOST /v1/formats/{handle}/{slug}/runs

What does a minimal request look like?

Send the brief as instruction with a generation_spend_cap_usd, which is required and has no default. The call answers 202 with an agent.run receipt. Take the result by webhook, or poll until next_action stops being poll_status.

curl -sS -X POST "https://api.sume.com/v1/agent/completions" \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: teaser-123-v1" \
  -d '{
    "instruction": "Make a 15-second vertical 9:16 promo for https://shop.example.com/p/123 with voiceover and captions. Deliver one MP4.",
    "generation_spend_cap_usd": 10,
    "communication": { "webhook_url": "https://example.com/hooks/sume" }
  }'

# Poll the run until next_action is no longer poll_status
curl -sS "https://api.sume.com/v1/agent-runs/$RUN_ID" \
  -H "Authorization: Bearer $SUME_API_KEY"

Should I use Agent Completions or a Format?

Sume's docs put it this way: use Agent Completions when the task itself varies per call, and a Format when you have a saved workflow and only the inputs change. Both run the same agent and return the same receipt shape.

Run the Sume video agent from your backend covers every Agent Completions field. What is a Sume Format? covers saving a recipe and calling it. What is a video agent? defines the term, and How to make a video ad with AI walks the chat path where a person approves a draft first.

Which other vendors document a video agent API?

Answers to this question often name HeyGen, Runway, and Synthesia. This is what each one's own developer docs said on 2026-09-28:

  • HeyGen: its Prompt to Video docs describe POST /v3/video-agents. You send a prompt of 1–10,000 characters, and "the agent handles scripting, avatar selection, scene composition, and rendering." The call returns a session_id; the finished video carries a video_url, and the response fields include captioned_video_url and subtitle_url. Sume vs HeyGen compares the two products.
  • Synthesia: its Create a video endpoint takes input, "an array of objects that each describe a clip of a multi-clip video," each naming an avatar, with the script as scriptText or as uploaded scriptAudio. You write the script per scene, as in the avatar row above.
  • Runway: its developer docs index says Runway Dev "exposes generative video, image, and audio models over HTTP"; most generation endpoints return a task id you poll. The same index lists Recipes, "prebuilt multi-step video and image workflows such as product ads and ad localization."

What are the limits?

The Agent Completions, Runs and results, and Errors and spend pages set these bounds:

  • A spend cap on every run. Agent Completions has no default for generation_spend_cap_usd; omitting it is 400 invalid_request. A Format run's cap goes up to the platform maximum of $500; a run that names no cap inherits the Format's, which is $400 if the Format never set one.
  • Minutes, not seconds. Format runs that make video take minutes, and the docs say long-form host video typically finishes in 15 to 30 minutes. A Format run still going 90 minutes after created_at is force-finalized as failed, sooner if it is older than 25 minutes and has been silent for 10.
  • No person in the loop. A Format run does not stop to ask anyone. A gate it cannot pass without a person, such as no avatar matching the brief, fails with unattended_blocked.
  • Not a chat stream. Agent Completions returns a run receipt, not choices[]. Streaming, continuing a prior thread, and non-image attachments are not available yet; up to 30 images can ride along.
  • Metered cost. Generation draws on the rates on API pricing, plus a 5.5% agent fee by default, and the receipt's usage records what the run spent.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume