Time a 360p Omni draft vs 720p: separate queue wait from run time
Google says the 360p draft is up to 60% faster. Measure it on Sume with a Python script that splits pending time from in_progress time.

To check Google's claim that the Omni 1.1 Flash 360p draft is up to 60% faster than 720p, time the stage that matters: how long the job sits in in_progress, not how long it waits in pending. On Sume a job can queue behind other jobs of your workspace, so wall-clock from submit mixes the two. The script below records both.
The claim and what it covers
The announcement says the 360p draft mode runs up to 60% faster and at a third of the cost of the standard 720p mode (Google, read 2026-10-05). "Up to" is a ceiling, and Google gives no clip length. Your speedup will be a number you measure for your clip length and your load.
Why queue time must be split out
Sume admits paid generation queue-first. A valid job can be accepted as queued when your workspace is at its concurrency limit, and it moves to processing when a slot opens (Sume docs: Generation admission, read 2026-10-05). On the Free plan the processing concurrency is 1; Pro is 4. Run the 360p and 720p tests one after the other on a quiet workspace, or the second job can queue behind the first and look slower for the wrong reason.
The /v1/videos poll statuses are pending, in_progress, completed, failed and cancelled (Sume docs: Video generation, read 2026-10-05). The timestamp where the status first reads in_progress is the start of the run.
Script
Each run bills like any job, so keep it to a few clips. The polling step is 3 seconds, so the result has about that much error.
import os, time, requests
URL = "https://api.sume.com/v1/videos"
AUTH = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def run(resolution):
body = {"model": "gemini-omni-flash-1.1",
"prompt": "A hand pours coffee into a white cup, close-up",
"resolution": resolution, "duration": 6, "aspect_ratio": "16:9"}
job = requests.post(URL, json=body, headers=AUTH).json()
t0, started = time.time(), None
while True:
s = requests.get(job["polling_url"], headers=AUTH).json()
now = time.time()
if s["status"] == "in_progress" and started is None:
started = now
if s["status"] in ("completed", "failed", "cancelled"):
break
time.sleep(3)
started = started or now
return {"res": resolution, "status": s["status"],
"queued_s": round(started - t0, 1), "run_s": round(now - started, 1)}
for res in ["360p", "720p", "360p", "720p"]:
print(run(res))How to read it
If the speedup you see is smaller than 60%, that does not contradict Google's wording. It means your clip and load are not the best case. The price ratio is separate and is checkable from the catalog.
- Compare
run_sfor the same clip length, resolution by resolution. Ignorequeued_s. - Alternate the tiers, as the loop does, so a slow provider period hits both.
- If a run shows
queued_sof more than a few seconds, your concurrency slot was busy; wait and rerun it. - Four runs is a spot check. A rollout decision needs more clips and more hours.
Report the number in the right units
State the result as seconds of in_progress per output second, so a 6-second clip that runs 40 seconds reads as 6.7 s per output second. Then you can compare it with a 10-second clip, and you can check whether the 360p gain grows with length. Keep the per-second price next to it, because speed and cost together decide whether a draft is worth a slot.
Cost is simple. Sume bills Omni per output second by resolution, so a 6-second 360p job is 6 times the 360p rate, and the same clip at 720p is 6 times the 720p rate. The catalog and the pricing page hold the live numbers.
Run each resolution at least three times and compare medians, not single runs. Queue time and generation time move independently, so one slow run says little about the tier.
Sources
Related posts
More in Developers
- Timeline refuses a 180-second Short: clips end more than 0.5 s early
Timeline 1.0 allows video coverage to stop at most 0.5 seconds before the end of the audio spine. Check slot ends in Python before you plan a 180-second Short.
- Timeline plan first: the unbilled estimate for a 30-second Short
POST /v1/timeline-1.0/plan compiles your cut without a job or a charge and returns duration, segments and estimated cost. A 30-second render bills one minute.
- Timeline plan for a 40-second Omni stitch: four segments, one minute
Run Sume's unbilled timeline plan on four 10-second Omni clips to see segment_count, billable_minutes and the cost before you render. Request body included.
- A 30-line Node proxy so a browser can start a Sume video, no key
A node:http server with only POST /render and GET /status/:id. It fixes the model and clip size and keeps SUME_API_KEY on the server, away from the browser.
Written by Sume