AI video generator from text: the three fields that set the price
In a text-to-video request on Sume, model, resolution and duration set the price. Prompt length and aspect ratio do not. Worked Wan 3.0 and Omni numbers.

A text-to-video request on Sume needs only model and prompt. Of the other fields, three move the bill: the model you pick, resolution, and duration. For per-second models the price is the catalog list rate times 1.25 times the seconds, rounded up to the cent. A longer prompt does not cost more.
The minimum body
These are the fields the docs mark as required, plus the three that decide cost:
{
"model": "wan-3.0",
"prompt": "A slow dolly shot across a desk with a notebook and a lamp",
"resolution": "720p",
"duration": 6,
"aspect_ratio": "16:9"
}What each field does to the bill
Docs say a higher resolution takes more time and has a higher price. For Wan 3.0 the catalog lists three per-second rates, so one six-second prompt can bill three different amounts:
| Resolution | Sume rate per second | 6-second clip |
|---|---|---|
| 480p | $0.0625 | $0.38 |
| 720p | $0.125 | $0.75 |
| 1080p | $0.250 | $1.50 |
Fields that do not move it
aspect_ratio is a shape, not a price tier. prompt has no per-character rate in the video catalog. callback_url only changes how you learn the job finished. Sume's docs also state that size, seed and provider.options return 400 unsupported_parameter on /v1/videos, so those fields never reach the provider and never price anything.
When duration is the cost lever
Duration scales linearly on per-second models, so halving a clip halves the rate-based price before rounding. The docs list a floor and ceiling per model: Wan 3.0 accepts 2 to 30 seconds, and Gemini Omni Flash 1.1 accepts 3 to 10. A 3-second Omni clip at 720p bills $0.38, while a 10-second one bills $1.25.
Seedance models are priced per video token rather than per second, so read their row in GET /v1/video-router/models before you estimate one.
A worked example with a choice at each field
Say you need a 6-second vertical clip. Pin Wan 3.0 at 480p and it bills $0.38. Move to 720p and the same prompt bills $0.75. Switch the model to Gemini Omni Flash 1.1 at 720p and it bills $0.75. The prompt is identical in all three. Only the fields in the table above changed.
That is why a cost review of a video app should look at defaults, not at prompts. If your UI sends resolution: "1080p" by default, every clip is billed at the top tier. The Auto create controls on /v1/videos default to 720p and 8 seconds, so an Auto request with no resolution or duration lands on a mid tier.
After a job completes, the poll response includes usage.cost. Compare it with your estimate once per batch. If they differ by more than a cent per job, check which resolution and duration reached Sume. Never assume that a missing field means the cheapest tier, because the model fills in its own default.
Before you ship anything, read the live pages again: the catalog is public, the pricing page is public, and the docs describe the request fields. A blog post is a snapshot. The catalog, the plan grid and the error table are the things that change, so write your code to read them instead of copying numbers from a page, and re-check when a new model is added.
A good habit is a small log line per submit with the model, resolution, duration, estimated cost, job id and the Idempotency-Key you used. When a job misbehaves, those six fields answer most of the questions support will ask, and they let you compare your estimate with usage.cost and the usage ledger without re-running anything.
Sources
Related posts
More in Developers
- AI video generator with no sign-up: what Sume needs before clip one
Sume cannot render video without an account, an API key and a funded balance. Four public reads work with no key, so you can check models first.
- AI video seed on Sume: no model accepts one, lock shots with frames
Every Sume v1 video model reports seed false and rejects the field. To repeat a look, lock first and last frames, references, resolution and a stable prompt.
- Write Amazon's contains-synthetic-performer tag with exiftool
Amazon wants contains-synthetic-performer in XMP dc:subject on photorealistic AI people. Sume does not write it; tag the downloaded file yourself with exiftool.
- Nova 2 Sonic cut hallucinations 88%: run a read-back check on Sume
AWS says its May Nova 2 Sonic refresh cut hallucinations 88% on an internal set. Test any voice on your own text: Sume TTS, then STT, then a diff.
Written by Sume