AI video generator for YouTube: Sume models that do 16:9 at 1080p

Four Sume video models accept 16:9 at 1080p. Their per-second rates and the cost of a 5-second and 8-second B-roll clip at 1080p.

5 min readSume
All posts

For a landscape YouTube clip, ask the catalog for 16:9 at 1080p. On Sume, Seedance 2.5 and 2.0, Wan 3.0, Gemini Omni Flash 1.1 and MiniMax H3 Max all list 1080p and take 16:9. MiniMax H3 stops at 768p natively, so it is out. This page is about generating the footage; for channel upload rules, check YouTube's own help pages.

The 1080p rates

Three of the four are priced per second, so you can compare them directly. Seedance bills per video token, so read its row in the catalog.

1080p per-second price and 5 s / 8 s clip price, list x 1.25 rounded up (catalog rates, read 2026-10-05)
ModelPer second5 s clip8 s clipDuration range
gemini-omni-flash-1.1$0.1875$0.94$1.503-10 s
minimax-h3-max$0.200$1.00$1.605-15 s
wan-3.0$0.250$1.25$2.002-30 s

Request for a 16:9 1080p clip

aspect_ratio and resolution go in the same body. Replace the model with any row above:

curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: yt-broll-001" \
  -d '{
    "model": "wan-3.0",
    "prompt": "Aerial shot of a coastal road at golden hour",
    "resolution": "1080p",
    "aspect_ratio": "16:9",
    "duration": 8
  }'

Picking among the three

Omni Flash 1.1 is the cheapest per second at 1080p and always renders synced audio, but it caps at 10 seconds. H3 Max accepts 5 to 15 seconds and renders native stereo audio. Wan 3.0 is the only one of the three that reaches 30 seconds and the only one with a 2-second floor, at the highest per-second rate. For a ten-second B-roll insert, Omni 1.1 Flash is the low-cost fit; for a 20-second establishing shot, Wan 3.0 or Seedance 2.5 are the models whose range reaches it.

Because these are catalog facts and not a guarantee, call GET /v1/videos/models before you pin: the docs note that some catalog models only appear when their provider is configured.

A cheaper route: draft at a low tier

A 1080p render is the final step, not the first. For a channel workflow, render the shot at a low tier, approve it, and then re-render at 1080p with the same prompt. Omni at 360p is $0.0375 a second, against $0.1875 at 1080p, so a 5-second draft is $0.19 against $0.94 for the final. Re-rendering is a new job with a new price, so count both in the budget.

One caution: a re-render with the same prompt is not guaranteed to match the draft. The docs say no v1 model accepts seed, so you cannot pin the output. Treat the draft as a look at composition and motion, not as an exact preview.

Aspect ratio for the Shorts shelf

Everything above is 16:9. If the same clip is also for a vertical placement, send a second request with aspect_ratio: "9:16". Omni and Wan both list it. The price depends on resolution and duration, not on aspect ratio, so the vertical version costs the same as the landscape one at the same settings. Check YouTube's own help pages for the current specs before you set the final size.

Before you ship anything, read the live pages again: the catalog is public, the pricing page is public, and the docs describe the request fields. A blog post is a snapshot. The catalog, the plan grid and the error table are the things that change, so write your code to read them instead of copying numbers from a page, and re-check when a new model is added.

A good habit is a small log line per submit with the model, resolution, duration, estimated cost, job id and the Idempotency-Key you used. When a job misbehaves, those six fields answer most of the questions support will ask, and they let you compare your estimate with usage.cost and the usage ledger without re-running anything.

Sources

Related posts

More in Models

All Models posts

Written by Sume