Text to video API: how it works, what it returns, cost
A text to video API takes a prompt and returns a job id, not a video: poll it or take a webhook, then download the file. Fields, flow and prices.

A text-to-video API takes a written prompt, plus optional settings such as length, resolution and aspect ratio, and returns a job id rather than a finished video. Generation takes long enough that you poll the job, or wait for a webhook, until it completes, and then download the video file.
On Sume that request is POST /v1/videos, with a catalog model id such as seedance-2.5, or sume/auto to let Sume pick. The fields and flow below come from the Video generation docs and the Sume API reference, read on 2026-09-28; "current code" marks behavior read from the API code.
What goes in a text-to-video request?
Two fields are required; the rest narrow the output. Each model lists the values it accepts in GET /v1/videos/models, and in current code a value it doesn't list is refused with 400 unsupported_capability.
| Field | Required | What it sets |
|---|---|---|
model | Yes | A catalog id, or sume/auto |
prompt | Yes | The description of the video |
duration | No | Length in whole seconds |
resolution | No | For example 720p or 1080p |
aspect_ratio | No | For example 16:9 or 9:16 |
generate_audio | No | An audio track; defaults to the model's audio capability |
callback_url | No | An HTTPS URL to notify when the job finishes |
How do I call a text-to-video API?
Submit, poll, download. The submit answers 202 Accepted with a job id, a polling_url and status: pending. Poll until status is completed; the docs suggest about 30 seconds between polls, and a job typically takes 30 seconds to several minutes. The download route needs the same API key and redirects to the file, so curl needs -L. Text to video API in Python shows the same loop in Python.
# 1. Submit (a retry with the same key and body returns the same job)
curl -X POST "https://api.sume.com/v1/videos" \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: t2v-beach-001" \
-d '{
"model": "seedance-2.5",
"prompt": "A golden retriever playing fetch on a sunny beach",
"duration": 8,
"resolution": "720p",
"aspect_ratio": "16:9"
}'
# 2. Poll until "status" is "completed"
curl "https://api.sume.com/v1/videos/job_123" \
-H "Authorization: Bearer $SUME_API_KEY"
# 3. Download the video
curl -L "https://api.sume.com/v1/videos/job_123/content?index=0" \
-H "Authorization: Bearer $SUME_API_KEY" --output video.mp4What statuses can a text-to-video job have?
Five: pending (queued), in_progress, completed, failed (read the error field), and cancelled. Instead of polling, send a callback_url: Sume POSTs a signed webhook once the job reaches a terminal state. Poll AI video job status covers the loop.
Which models can make video from text alone?
On Sume: seedance-2.5, seedance-2-mini, seedance-2, seedance-2-fast, kling-3, wan-3.0, minimax-h3, minimax-h3-max, gemini-omni-flash-1.1. grok-imagine-video-1.5 needs a first-frame image. Lengths differ: seedance-2.5 and wan-3.0 reach 30 seconds and every other model tops out at 15. With sume/auto, Sume picks the family and never discloses which family ran.
How much does a text-to-video API cost?
Sume bills each clip per second or per 1,000 video tokens, depending on the model, at the provider's list price × 1.25, reserved when you submit; the finished job reports the billed amount in usage.cost. Ten seconds reserves $0.75 on minimax-h3 at 768p, $1.25 on wan-3.0 at 720p, $1.40 on kling-3 without audio, $5.78 on seedance-2.5 at 720p 16:9, plus a 5.5% agent fee by default. How much does an AI video cost? compares every model.
What are the limits of a text-to-video call on Sume?
A single text-to-video call on Sume can't:
- Make a long video in one call: one clip is 30 seconds at most. Chain clips for more, as AI video length limits by model explains.
- Take exact pixel sizes or a seed: no v1 model accepts
sizeorseed, and both return400 unsupported_parameter. - Show your real product or person: from text, the model invents every frame. Start from a picture instead, as in Text to video and image to video.
Sources
Related posts
More in Developers
- Webhook security best practices: a receiver checklist
Accept HTTPS only, verify an HMAC over the raw body in constant time, reject stale timestamps, dedupe on the event id, and answer 2xx fast.
- Webhook vs API: what's the difference?
An API call is your code asking a server for something; a webhook is the server calling your URL when something happens. The two work together.
- What is a dead letter queue? DLQs for AI job pipelines
A dead-letter queue holds messages that failed processing too many times, so they stop looping and can be inspected. Which AI job failures go there.
- What is a video API? The five kinds, explained
A video API lets code make, edit or deliver video over HTTP. The five kinds, what each one takes and returns, and how to tell which one you need.
Written by Sume