A 60-second Short as twenty 3-second Omni cuts: $7.50 at 720p
Twenty 3-second Omni Flash 1.1 clips at 720p cost $7.50 and at 1080p $11.25, the same as six 10 s clips. Twenty slots pass 12, so Timeline chunks the render.

Gemini Omni Flash 1.1 accepts as little as 3 seconds a job, so a fast-cut 60-second vertical Short can be twenty separate 3-second clips. They cost $7.50 at 720p and $11.25 at 1080p (the rate is linear, so six 10-second clips cost the same), plus $0.10 for the one-minute Timeline render.
Fast cuts suit Shorts, which are vertical and up to 3 minutes (YouTube Help, read 2026-10-09). The catch is operational: twenty jobs, twenty slots and a render that Timeline will chunk.
Cost
Omni Flash lists 360p, 720p, 1080p and 4K in 16:9 or 9:16, 3 to 10 seconds a job.
| Resolution | Rate per second | One 3 s cut | Twenty cuts | With Timeline ($0.10) |
|---|---|---|---|---|
| 720p | $0.125 | $0.38 | $7.50 | $7.60 |
| 1080p | $0.1875 | $0.5625 | $11.25 | $11.35 |
Building the body
Timeline takes 1 to 200 slots. With render.strategy auto it chunks past 12 segments, and it refuses single above 12 (render_strategy_unsafe). More than 8 adjacent fades fail with too_many_chained_transitions, so the sample puts a fade on every third join and hard-cuts the rest.
import json
slots = []
for i in range(20):
slot = {"source_url": f"https://media.sume.com/artifacts/artf_demo/cut{i + 1:02d}.mp4",
"start": i * 3, "duration": 3}
if i and i % 3 == 0: # a fade on every third join, hard cuts between
slot["transition"] = {"type": "fade", "duration": 0.25}
slots.append(slot)
body = {"audio": {"mode": "silence", "duration_seconds": 60},
"video": slots, "render": {"strategy": "auto"},
"output": {"width": 1080, "height": 1920}}
print(json.dumps(body))Limits
- Twenty prompts means twenty places for the style to drift; pass the same reference image to every request.
- A cut shorter than 3 seconds is not possible from Omni; use Timeline's duration (min 0.2 s) to shorten a clip inside the render instead.
- Plan first with POST /v1/timeline-1.0/plan, which is unbilled, to see segment_count and billable_minutes.
- Each slot start must increase and the first must be 0.
Sources
Related posts
More in Use cases
- Spelling bee: 200 word-pronunciation clips on Sume TTS, one job each
200 words of about 12 characters each cost $0.1140 on Sume TTS as 200 jobs, $0.000570 each. One 2,400-character job is $0.1140 but gives one file.
- Split a 14-minute episode into 20 clips with overlaps: 2 cents
Detach the audio once ($0.01), then one timeline audio split with 20 overlapping ranges ($0.01). A 860-second track fits the 900-second detach cap.
- Sponsored Snaps headline: 24 to 28 characters, check a batch
Snapchat's Sponsored Snaps page recommends a 24-28 character headline on a 9:16 asset. Check a batch of headlines in Python before you queue the Sume clips.
- Square 1:1 AI video for YouTube Shorts: Seedance 2.0 15 s from $2.64
YouTube counts square or vertical uploads up to 3 minutes as Shorts. A 15-second 1:1 Seedance 2.0 clip on Sume is $2.64 at 480p, $5.67 at 720p, $12.76 at 1080p.
Written by Sume