What is a video API? The five kinds, explained
A video API lets code make, edit or deliver video over HTTP. The five kinds, what each one takes and returns, and how to tell which one you need.

A video API is an HTTP interface that lets your code make, edit or deliver video instead of doing it by hand in an app. The name covers five different products: hosting and streaming APIs, rendering APIs (a JSON document in, an MP4 out), generative model APIs (a prompt in, new footage out), avatar APIs (a script in, a talking presenter out), and agent APIs that turn a brief into a finished video.
The category definitions below are plain descriptions, not any one vendor's product. The Sume examples come from Sume basics and the model guides it links, read on 2026-09-28. Sume's docs cover the last four kinds and describe no hosting or streaming product.
What are the kinds of video API?
Ask what goes in and what comes out. That one question separates the five kinds:
| Kind | You send | You get back | Sume example |
|---|---|---|---|
| Hosting and streaming | A finished video file | Storage, playback URLs, a player | None. Sume returns generated outputs as media.sume.com URLs, its public artifact contract |
| Rendering (JSON to video) | A document that orders clips and audio | One rendered file | Timeline 1.0: one audio spine plus ordered video[] slots in, one MP4 out |
| Generative model | A prompt, optionally images | New footage | POST /v1/videos, from text prompts and optional reference images |
| Avatar | A script and a presenter | A talking-presenter clip | Avatar 1.0: create a reusable avatar, then POST /v1/avatar-1.0/talking-video |
| Agent | A brief or a saved recipe's inputs | Finished artifacts | A Format run, or POST /v1/agent/completions for a one-off task |
Is a JSON-to-video API the same as an AI video API?
No. A rendering API assembles footage you already have: it cuts, orders and joins clips to a timeline you describe in JSON. It never invents a shot. A generative API makes footage that did not exist before, from a prompt. Many products need both: generate the clips, then render them into one video.
Sume's Timeline 1.0 shows the rendering side. It takes a declarative document and returns one MP4, and callers never send filtergraphs, codecs or shell fragments. A render needs audio.duration_seconds from 1 to 1800 and 1 to 200 video[] slots, and every URL must already be the workspace's own media.sume.com artifact or asset, such as an earlier Sume job's output. How to assemble a long-form video walks through a render.
How does a generative video API call work?
Sume's video generation API is asynchronous: you submit, get a job id, and check back. POST /v1/videos answers 202 Accepted with the job's id, a polling_url and status: "pending". You poll that URL until the job completes, then download the file.
The model field takes a catalog id such as seedance-2.5, or sume/auto to let Sume pick. Text-to-video API covers the request fields and prices.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "sume/auto", "prompt": "A slow pan across a desk at sunrise"}'What do avatar and agent video APIs add?
An avatar API keeps a presenter as a reusable resource. Sume's Avatar 1.0 is a two-step workflow: create an avatar once, then use it to generate talking videos from scripts or multi-scene inputs.
An agent API sits one level higher: it decides which tools to call and in what order, then returns the finished result. Video agent API covers what Sume's agent calls take and return, and compares them with the model and render routes for one brief.
Which kind of video API do I need?
- You already have finished files and need to store, play or stream them: a hosting or streaming API. Sume is not one.
- You have clips and audio and need them joined in a fixed order: a rendering API, such as Timeline 1.0.
- You need new footage from a prompt or a still image: a generative model API, such as
POST /v1/videos. - You need the same presenter speaking many scripts: an avatar API.
- You need a finished, edited deliverable from a brief, and don't want to write the orchestration yourself: an agent API, such as a Format run. How to evaluate an AI video generation API has a checklist for comparing vendors once you know the kind.
Sources
Related posts
More in Developers
- Edit decision list (EDL): what it is, with an example
An edit decision list (EDL) is the ordered list of edits that rebuilds a cut: source, track, transition, and timecodes. An example and its JSON form.
- What is serverless inference? How it works and is billed
Serverless inference means calling a hosted AI model over HTTP with no servers of your own. Providers bill for compute time or for each output.
- yuv420p10le vs yuv420p: 8-bit vs 10-bit pixel formats
yuv420p and yuv420p10le are both planar YUV 4:2:0. yuv420p stores 8 bits per sample; yuv420p10le stores 10, little-endian, in 16-bit words.
- Where to store API keys: server, CI, and local dev
Keep API keys server-side: an env var fed by a secret manager in production, your CI's secret store in pipelines, a git-ignored file on your laptop.
Written by Sume