Text to video API: how it works, what it returns, cost

A text to video API takes a prompt and returns a job id, not a video: poll it or take a webhook, then download the file. Fields, flow and prices.

5 min readSume
All posts

A text-to-video API takes a written prompt, plus optional settings such as length, resolution and aspect ratio, and returns a job id rather than a finished video. Generation takes long enough that you poll the job, or wait for a webhook, until it completes, and then download the video file.

On Sume that request is POST /v1/videos, with a catalog model id such as seedance-2.5, or sume/auto to let Sume pick. The fields and flow below come from the Video generation docs and the Sume API reference, read on 2026-09-28; "current code" marks behavior read from the API code.

What goes in a text-to-video request?

Two fields are required; the rest narrow the output. Each model lists the values it accepts in GET /v1/videos/models, and in current code a value it doesn't list is refused with 400 unsupported_capability.

From Video generation, read 2026-09-28.
FieldRequiredWhat it sets
modelYesA catalog id, or sume/auto
promptYesThe description of the video
durationNoLength in whole seconds
resolutionNoFor example 720p or 1080p
aspect_ratioNoFor example 16:9 or 9:16
generate_audioNoAn audio track; defaults to the model's audio capability
callback_urlNoAn HTTPS URL to notify when the job finishes

How do I call a text-to-video API?

Submit, poll, download. The submit answers 202 Accepted with a job id, a polling_url and status: pending. Poll until status is completed; the docs suggest about 30 seconds between polls, and a job typically takes 30 seconds to several minutes. The download route needs the same API key and redirects to the file, so curl needs -L. Text to video API in Python shows the same loop in Python.

# 1. Submit (a retry with the same key and body returns the same job)
curl -X POST "https://api.sume.com/v1/videos" \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: t2v-beach-001" \
  -d '{
    "model": "seedance-2.5",
    "prompt": "A golden retriever playing fetch on a sunny beach",
    "duration": 8,
    "resolution": "720p",
    "aspect_ratio": "16:9"
  }'

# 2. Poll until "status" is "completed"
curl "https://api.sume.com/v1/videos/job_123" \
  -H "Authorization: Bearer $SUME_API_KEY"

# 3. Download the video
curl -L "https://api.sume.com/v1/videos/job_123/content?index=0" \
  -H "Authorization: Bearer $SUME_API_KEY" --output video.mp4

What statuses can a text-to-video job have?

Five: pending (queued), in_progress, completed, failed (read the error field), and cancelled. Instead of polling, send a callback_url: Sume POSTs a signed webhook once the job reaches a terminal state. Poll AI video job status covers the loop.

Which models can make video from text alone?

On Sume: seedance-2.5, seedance-2-mini, seedance-2, seedance-2-fast, kling-3, wan-3.0, minimax-h3, minimax-h3-max, gemini-omni-flash-1.1. grok-imagine-video-1.5 needs a first-frame image. Lengths differ: seedance-2.5 and wan-3.0 reach 30 seconds and every other model tops out at 15. With sume/auto, Sume picks the family and never discloses which family ran.

How much does a text-to-video API cost?

Sume bills each clip per second or per 1,000 video tokens, depending on the model, at the provider's list price × 1.25, reserved when you submit; the finished job reports the billed amount in usage.cost. Ten seconds reserves $0.75 on minimax-h3 at 768p, $1.25 on wan-3.0 at 720p, $1.40 on kling-3 without audio, $5.78 on seedance-2.5 at 720p 16:9, plus a 5.5% agent fee by default. How much does an AI video cost? compares every model.

What are the limits of a text-to-video call on Sume?

A single text-to-video call on Sume can't:

  • Make a long video in one call: one clip is 30 seconds at most. Chain clips for more, as AI video length limits by model explains.
  • Take exact pixel sizes or a seed: no v1 model accepts size or seed, and both return 400 unsupported_parameter.
  • Show your real product or person: from text, the model invents every frame. Start from a picture instead, as in Text to video and image to video.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume