How to build your own AI video generator (no training)

Build your own AI video generator from a front end, a small backend and a video model API. You don't train a model; your backend holds the key.

5 min readSume
All posts

To build your own AI video generator, you don't train a model: you build a front end and a small backend, and connect them to a hosted video model API. The backend holds the API key, submits each generation job, learns when it finishes from a webhook or by polling, and hands the user the finished file.

The Sume details below come from its video generation, authentication and webhooks docs, read 2026-09-28. The same architecture works with any video API that runs jobs asynchronously.

What are the parts of an AI video generator?

Five parts, and only one of them is AI. Everything else is ordinary web plumbing you already know how to build.

From Sume's Authentication, Video generation, Webhooks and Jobs and results docs, read 2026-09-28.
PartWhat it doesRule from the docs
Front end (site or app)Takes the prompt, shows status, plays the resultCalls your backend; never holds the API key
Backend endpointChecks the user may generate, builds the request, attaches the keyValidate input and enforce your own authorization first
Video model APIRuns the model and returns a job id at oncePOST /v1/videos, then poll or take a webhook
Webhook receiverHears when a job endsPublic HTTPS URL; terminal events only
Job tableMaps each job id to the user who askedStore the job id so you can recover after restarts

How does the backend submit a job?

The browser posts the prompt to your endpoint. The endpoint checks the user, calls the video API with the key from its environment, and saves the returned job id against that user. Derive the Idempotency-Key from something stable, such as a request id the front end generates once per click, so a double-click or a network retry returns the original job instead of a second paid one. A replay needs the same key and the same body.

// Server only. The key never reaches the browser.
export async function POST(request: Request) {
  const user = await requireUser(request); // your own auth
  const { prompt, requestId } = await request.json();
  const res = await fetch("https://api.sume.com/v1/videos", {
    method: "POST",
    headers: {
      Authorization: `Bearer ${process.env.SUME_API_KEY}`,
      "Content-Type": "application/json",
      "Idempotency-Key": `${user.id}:${requestId}`,
    },
    body: JSON.stringify({
      model: "seedance-2",
      prompt,
      resolution: "720p",
      callback_url: "https://example.com/hooks/sume",
    }),
  });
  const job = await res.json();
  if (!res.ok) return Response.json(job, { status: res.status });
  await saveJob({ jobId: job.id, userId: user.id }); // your database
  return Response.json({ jobId: job.id });
}

How does the user get the finished video?

Video takes a while: Sume's docs say generation "typically takes 30 seconds to several minutes". With callback_url set, Sume POSTs a signed job.completed, job.failed or job.canceled event to your receiver once the job is terminal, making up to 10 delivery attempts. Deliveries can repeat, so treat job_id as your idempotency key, and keep polling as a backup, because delivery is "never the only recovery path" (Webhooks).

When the job completes, look up the user by job_id and give the front end the file. Hand it the public media.sume.com URL in the webhook's payload.artifacts, not the unsigned_urls entry, which needs your key; Download a generated video shows both. Anyone who has a media.sume.com URL can fetch it, so copy the file into your own storage if users must not see each other's videos.

Should I call one model or a whole workflow?

Sume's docs put the rule in one line: "If you only need a single clip or image, call the model; if you need a packaged workflow, call a Format" (The basics). A generator that turns a prompt into one clip calls a model. A generator that writes a script, adds a voiceover, B-roll and captions, and assembles them is a workflow; run it as a saved Format. How to embed AI video generation in your product covers that path, with per-customer keys, spend caps and webhooks.

What limits should I plan for?

  • Concurrency: your plan caps how many jobs run at once; extra jobs wait as queued. Video job concurrency and queueing has the numbers.
  • Cost: every generation spends your balance, reserved on submit at provider list × 1.25 (Video generation). Add your own per-user limits, since Sume authenticates you, not your users.
  • Webhook URLs: public HTTPS only; localhost, private-network and non-HTTPS URLs are rejected (Webhooks).
  • Hardware: none. The provider runs the model; Do you need a GPU for AI video? explains when you would.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume