MiniMax Video Agent template API: submit, poll, download

MiniMax Video Agent builds a video from a template_id plus your media and text. The endpoints, statuses, and what Sume offers for template-driven video.

5 min readSume
All posts

MiniMax's Video Agent turns a template_id plus your images, videos and text into a finished video through three steps: submit to POST /v1/video_template_generation, poll GET /v1/query/video_template_generation, then download from the returned video_url. Sume has no catalog of public templates on the video endpoint; its nearest concept is a Format, a saved recipe you call by name.

The MiniMax side comes from the template guide and video generation guide, read on 2026-10-02. Sume's side is from the Format API and Video generation.

How does the MiniMax template flow work?

A template defines placeholders. You fill media placeholders with images or videos and text placeholders with strings, then submit the task with the template id and your assets. The guide's worked example uses a template named 'Run for Life' with the id 393769180141805569, for stylized video, and points to a separate list of templates.

The task is asynchronous. Submit returns a task_id, and you poll with it until the task reports Success, which supplies video_url, or Fail, which supplies error details. Authentication is a bearer token in the Authorization header. The page gives no polling interval beyond asking for a reasonable one, and no price, so neither is claimed here.

MiniMax Video Agent calls, read 2026-10-02
StepMethod and URLResult
SubmitPOST https://api.minimax.io/v1/video_template_generationtask_id
CheckGET https://api.minimax.io/v1/query/video_template_generationSuccess with video_url, or Fail
DownloadFetch video_urlThe finished video

Is this the same as the H3 video API?

No. The H3 and H3 Max models run on the V2 create endpoint, with their own query path, and the video guide states durations of 4 to 15 seconds for H3 and 5 to 15 for H3 Max, with asset limits of 50 MB for video, 30 MB for images and 15 MB for audio. The template service is a separate product with its own paths and status words (Success and Fail rather than succeeded and failed).

If you already wrapped the H3 endpoints, do not reuse the same status parser for templates. Map each service's words to your own enum at the boundary, as you would for any second vendor.

What does Sume offer for template-driven video?

On POST /v1/videos there is no template id: you send a model, a prompt, optional frame images and references, and parameters such as duration and aspect_ratio. The model decides the look; nothing in the request selects a ready-made visual style by id.

The closer match is a Format: a saved production recipe with a house style and an output contract, which you call from your backend by name. A run gives you finished media on media.sume.com and, when you ask, JSON in a schema you supply. A Format is not a MiniMax template: you author or own it, and it does not fill placeholders from a marketplace list. If the template catalog is the reason you chose MiniMax, Sume does not replace it.

curl -sS "https://api.sume.com/v1/formats" \
  -H "Authorization: Bearer $SUME_API_KEY"

How do the polling contracts compare?

Both are submit, poll and fetch, but the words differ, so keep a mapping in code rather than in your head.

Status words, read 2026-10-02
StageMiniMax Video AgentSume job
RunningIn progress (not named)queued, processing
DoneSuccesscompleted
FailedFailfailed
StoppedNot stated for templatescanceled

Which should you use?

Pick MiniMax Video Agent if you want a fixed, designed look chosen by template id and you can accept its terms. Pick Sume if you want to choose among video models through one request shape, with idempotent retries and a job you can also read at /v1/jobs/{id}. Poll with backoff on either, and never resubmit a paid request just because your client timed out; see Jobs and results.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume