How to build your own AI video generator (no training)
Build your own AI video generator from a front end, a small backend and a video model API. You don't train a model; your backend holds the key.

To build your own AI video generator, you don't train a model: you build a front end and a small backend, and connect them to a hosted video model API. The backend holds the API key, submits each generation job, learns when it finishes from a webhook or by polling, and hands the user the finished file.
The Sume details below come from its video generation, authentication and webhooks docs, read 2026-09-28. The same architecture works with any video API that runs jobs asynchronously.
What are the parts of an AI video generator?
Five parts, and only one of them is AI. Everything else is ordinary web plumbing you already know how to build.
| Part | What it does | Rule from the docs |
|---|---|---|
| Front end (site or app) | Takes the prompt, shows status, plays the result | Calls your backend; never holds the API key |
| Backend endpoint | Checks the user may generate, builds the request, attaches the key | Validate input and enforce your own authorization first |
| Video model API | Runs the model and returns a job id at once | POST /v1/videos, then poll or take a webhook |
| Webhook receiver | Hears when a job ends | Public HTTPS URL; terminal events only |
| Job table | Maps each job id to the user who asked | Store the job id so you can recover after restarts |
How does the backend submit a job?
The browser posts the prompt to your endpoint. The endpoint checks the user, calls the video API with the key from its environment, and saves the returned job id against that user. Derive the Idempotency-Key from something stable, such as a request id the front end generates once per click, so a double-click or a network retry returns the original job instead of a second paid one. A replay needs the same key and the same body.
// Server only. The key never reaches the browser.
export async function POST(request: Request) {
const user = await requireUser(request); // your own auth
const { prompt, requestId } = await request.json();
const res = await fetch("https://api.sume.com/v1/videos", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.SUME_API_KEY}`,
"Content-Type": "application/json",
"Idempotency-Key": `${user.id}:${requestId}`,
},
body: JSON.stringify({
model: "seedance-2",
prompt,
resolution: "720p",
callback_url: "https://example.com/hooks/sume",
}),
});
const job = await res.json();
if (!res.ok) return Response.json(job, { status: res.status });
await saveJob({ jobId: job.id, userId: user.id }); // your database
return Response.json({ jobId: job.id });
}How does the user get the finished video?
Video takes a while: Sume's docs say generation "typically takes 30 seconds to several minutes". With callback_url set, Sume POSTs a signed job.completed, job.failed or job.canceled event to your receiver once the job is terminal, making up to 10 delivery attempts. Deliveries can repeat, so treat job_id as your idempotency key, and keep polling as a backup, because delivery is "never the only recovery path" (Webhooks).
When the job completes, look up the user by job_id and give the front end the file. Hand it the public media.sume.com URL in the webhook's payload.artifacts, not the unsigned_urls entry, which needs your key; Download a generated video shows both. Anyone who has a media.sume.com URL can fetch it, so copy the file into your own storage if users must not see each other's videos.
Should I call one model or a whole workflow?
Sume's docs put the rule in one line: "If you only need a single clip or image, call the model; if you need a packaged workflow, call a Format" (The basics). A generator that turns a prompt into one clip calls a model. A generator that writes a script, adds a voiceover, B-roll and captions, and assembles them is a workflow; run it as a saved Format. How to embed AI video generation in your product covers that path, with per-customer keys, spend caps and webhooks.
What limits should I plan for?
- Concurrency: your plan caps how many jobs run at once; extra jobs wait as
queued. Video job concurrency and queueing has the numbers. - Cost: every generation spends your balance, reserved on submit at provider list × 1.25 (Video generation). Add your own per-user limits, since Sume authenticates you, not your users.
- Webhook URLs: public HTTPS only; localhost, private-network and non-HTTPS URLs are rejected (Webhooks).
- Hardware: none. The provider runs the model; Do you need a GPU for AI video? explains when you would.
Sources
Related posts
More in Developers
- How to get a Kling AI API key and authenticate requests
Create a Kling AI API key in the developer console, copy it once, and send it as a Bearer token. Legacy endpoints use an Access Key JWT instead.
- How to get a public URL for an image an API can fetch
Host the image where anyone can fetch it over HTTPS without a login: a public storage object, a public bucket URL, or your own site. Then test it.
- How to test an MCP server: Inspector, Postman, or curl
Test an MCP server with the MCP Inspector: connect, sign in or add an auth header, list tools, and call a read-only one. Postman and curl work too.
- HTTP 202 vs 201: created now, or accepted for later?
Return 201 Created when the resource exists by the time you respond, and 202 Accepted when the work will finish later. How 200, 201 and 202 differ.
Written by Sume