AI video generator from text: the three fields that set the price

In a text-to-video request on Sume, model, resolution and duration set the price. Prompt length and aspect ratio do not. Worked Wan 3.0 and Omni numbers.

5 min readSume
All posts

A text-to-video request on Sume needs only model and prompt. Of the other fields, three move the bill: the model you pick, resolution, and duration. For per-second models the price is the catalog list rate times 1.25 times the seconds, rounded up to the cent. A longer prompt does not cost more.

The minimum body

These are the fields the docs mark as required, plus the three that decide cost:

{
  "model": "wan-3.0",
  "prompt": "A slow dolly shot across a desk with a notebook and a lamp",
  "resolution": "720p",
  "duration": 6,
  "aspect_ratio": "16:9"
}

What each field does to the bill

Docs say a higher resolution takes more time and has a higher price. For Wan 3.0 the catalog lists three per-second rates, so one six-second prompt can bill three different amounts:

The same 6-second Wan 3.0 text prompt at three resolutions, list x 1.25 rounded up (catalog rates, read 2026-10-05)
ResolutionSume rate per second6-second clip
480p$0.0625$0.38
720p$0.125$0.75
1080p$0.250$1.50

Fields that do not move it

aspect_ratio is a shape, not a price tier. prompt has no per-character rate in the video catalog. callback_url only changes how you learn the job finished. Sume's docs also state that size, seed and provider.options return 400 unsupported_parameter on /v1/videos, so those fields never reach the provider and never price anything.

When duration is the cost lever

Duration scales linearly on per-second models, so halving a clip halves the rate-based price before rounding. The docs list a floor and ceiling per model: Wan 3.0 accepts 2 to 30 seconds, and Gemini Omni Flash 1.1 accepts 3 to 10. A 3-second Omni clip at 720p bills $0.38, while a 10-second one bills $1.25.

Seedance models are priced per video token rather than per second, so read their row in GET /v1/video-router/models before you estimate one.

A worked example with a choice at each field

Say you need a 6-second vertical clip. Pin Wan 3.0 at 480p and it bills $0.38. Move to 720p and the same prompt bills $0.75. Switch the model to Gemini Omni Flash 1.1 at 720p and it bills $0.75. The prompt is identical in all three. Only the fields in the table above changed.

That is why a cost review of a video app should look at defaults, not at prompts. If your UI sends resolution: "1080p" by default, every clip is billed at the top tier. The Auto create controls on /v1/videos default to 720p and 8 seconds, so an Auto request with no resolution or duration lands on a mid tier.

After a job completes, the poll response includes usage.cost. Compare it with your estimate once per batch. If they differ by more than a cent per job, check which resolution and duration reached Sume. Never assume that a missing field means the cheapest tier, because the model fills in its own default.

Before you ship anything, read the live pages again: the catalog is public, the pricing page is public, and the docs describe the request fields. A blog post is a snapshot. The catalog, the plan grid and the error table are the things that change, so write your code to read them instead of copying numbers from a page, and re-check when a new model is added.

A good habit is a small log line per submit with the model, resolution, duration, estimated cost, job id and the Idempotency-Key you used. When a job misbehaves, those six fields answer most of the questions support will ask, and they let you compare your estimate with usage.cost and the usage ledger without re-running anything.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume