Which AI video model for ads, product shots or talking heads?
On Sume, pick Wan 3.0 or Omni for ads, Kling 3 for silent product shots, and the Avatar Video route for talking heads. One table with rates and limits.

For an ad with sound, pick gemini-omni-flash-1.1 ($0.125 a second at 720p) or wan-3.0; for a silent product shot at 1080p, pick kling-3 ($0.14); for a talking head with lip-synced speech, use the Avatar Video route, POST /v1/avatar-1.0/talking-video. Rates here are list times 1.25 from the Video Router guide, read 2026-10-05.
Veo, Luma Ray3.2, Runway and HappyHorse are not in the Sume catalog, so the table covers only models you can call.
One-table pick
Sume bills the provider list price times 1.25 and rounds the job up to the cent, so a one-second figure here is a rate, not a charge.
| Job | Pick | Rate per second | Why |
|---|---|---|---|
| Short ad with sound | gemini-omni-flash-1.1 | $0.125 at 720p | Native audio, first and last frame, references |
| Long or multi-ratio ad | wan-3.0 | $0.0625 to $0.25 | 2 to 30 s, five ratios, references |
| Silent product shot | kling-3 | $0.14 at 1080p | Cheapest 1080p, first and last frame |
| Cheapest still-to-video | grok-imagine-video-1.5 | see usage.cost | Needs a first frame, 480p or 720p, 4 to 15 s |
| Talking head | Avatar Video route | see the avatar docs | Speech-driven face, not a general video row |
Not sure
Leave model as sume/auto and the router picks, with gemini-omni-flash-1.1 as the default. Pin a model only once a job has a clear need, such as a ratio or a clip longer than 10 seconds.
Before you pin, call GET /v1/videos/models and check supported_durations, supported_aspect_ratios and supported_frame_images.
- Sound needed, 10 s or less: Omni.
- Over 10 s: Wan or Seedance 2.5.
- Lip sync: Avatar Video.
Request
Let the router pick.
import os, time, requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
payload = {
"model": "sume/auto",
"prompt": "A founder holds the product up to camera in a bright studio",
"duration": 5,
"resolution": "720p",
}
job = requests.post("https://api.sume.com/v1/videos", headers=H, json=payload).json()
while True:
time.sleep(30)
s = requests.get(job["polling_url"], headers=H).json()
if s["status"] in ("completed", "failed", "cancelled"):
break
print(s["status"], s.get("unsigned_urls"), s.get("usage"))Sources
Related posts
More in Use cases
- Which LinkedIn ad formats can carry a Lead Gen Form?
LinkedIn lists single image, carousel, video, event, message, document and conversation ads for Lead Gen Forms. See which asset Sume can produce for each.
- Which live voice jobs fit Sume TTS? A four-case table
Sume TTS is submit-and-poll with no streaming. Fixed prompts and alert lines fit; a conversational turn does not. A four-case table, with costs.
- Which scripts can Sume burn into captions? Latin and Hangul, then test
Index-Translate covers 150 languages, but Sume's caption docs describe Latin and Hangul styles. Route by script and spend $0.20 on a test cue for the rest.
- Which SKUs get an AI video first? Rank by profit inside a fixed budget
Rank SKUs by 30-day gross margin and fill a fixed video budget at Sume prices: $0.625 Wan or $1.89 Seedance 2 clips plus $0.30 for captions and Timeline.
Written by Sume