Agent leaves model blank on generate_video: 720p 8 s costs $1.00

Omit payload.model and generate_video routes to sume/auto: 3 to 10 s, 720p and 8 s by default. At 720p the price runs from $0.38 for 3 s to $1.25 for 10 s.

4 min readSume
All posts

When an agent calls generate_video without payload.model, Sume routes to sume/auto. Auto's create controls default to 720p and 8 seconds, with 3 to 10 second clips at 16:9 or 9:16. At 720p, the billable price runs from $0.38 for 3 seconds to $1.25 for 10 seconds, and the 8-second default costs $1.00. The tools-and-gates page says to omit the model unless the user named a family.

The Auto defaults

Sume's Video Router page (read 2026-10-08) says Auto is the recommended path for new integrations and that Video 1.0 stays a compatibility alias for the same pipe. The Auto defaults are 720p and 8 seconds, with 3 to 10 second clips at 16:9 or 9:16. The default model behind Auto is Gemini Omni Flash 1.1, whose native synced audio is always on. The job still reports model: sume/auto, and Sume does not disclose which family ran.

Price by length at 720p

Sume bills the provider list price times 1.25, rounded up to the cent, as a function of the output seconds and the resolution. These are the billable amounts for the Gemini Omni Flash 1.1 catalog entry at 720p, and 1080p for comparison.

Gemini Omni Flash 1.1, billable cents per clip, as of 2026-10-08
Seconds720p1080p
338 cents57 cents
450 cents75 cents
563 cents94 cents
675 cents113 cents
8 (default)100 cents150 cents
10125 cents188 cents

What a blank model hides

Because the response says sume/auto, an agent cannot tell you which model produced the clip, and it cannot promise the same look next time. If the look matters, pin a catalog id from video-router_models instead. Pinning also changes the price: a 6-second 720p 9:16 clip on Seedance 2.5 is 347 cents, compared with 75 cents on the Auto default, about 4.6 times as much (347 divided by 75).

Three limits to put in the instruction so the agent does not fail a call:

  • Length must be 3 to 10 seconds on Auto; for 15 seconds, ask for two clips, or pin a model with a longer range.
  • Aspect ratio is 16:9 or 9:16, nothing else.
  • Audio is always on for this model; the API rejects generate_audio: false.

Estimate before the call

Ask the agent to call generate_video with dry_run=true and report the estimate before submitting. For ten 8-second clips the expected spend is 10 x $1.00 = $10.00, so a max_spend_usd of 10.5 leaves a small margin. Each clip still needs its own idempotency_key, and the wave can be waited on with one batch jobs_wait of up to 20 ids.

A ten-clip example end to end

Say the instruction is: make ten 8-second vertical clips for a product line. With Auto defaults, each is 100 cents, so ten is 1,000 cents, or $10.00. Pro takes all ten at once (24 accepted), while Free accepts only six at a time, so the agent would submit six and then the remaining four as slots free up.

Between submit and delivery, the agent needs one batch jobs_wait call per 55 seconds, and the job ids from the creates. If the first wave returns wait_slice_expired, it repeats the call. When all ten are completed, one jobs_result with job_ids reads the whole wave.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume