A Veo 3.1 call becomes a Sume Omni job in under 30 lines of Python
Replace a Veo 3.1 request with a Sume gemini-omni-flash-1.1 job: submit, poll every 30 seconds, download the mp4 and read usage.cost. Standard library only.

The smallest replacement for a Veo 3.1 preview call on Sume is one function that posts to https://api.sume.com/v1/videos with model: "gemini-omni-flash-1.1", polls the job until it is completed, downloads the first unsigned_urls entry and returns usage.cost. The script below does that with the Python standard library in under 30 lines. It is the shape the Sume docs describe: submit, receive a job id and polling URL, poll, download.
The script
Set SUME_API_KEY first. The script refuses to run when the key is empty. It sends an Idempotency-Key so that a retried submit returns the original job, which Sume documents for this route. It polls every 30 seconds, the interval the docs suggest, and it stops on failed or cancelled. The last line runs one 8-second 720p clip.
import json, os, time, urllib.request
BASE = "https://api.sume.com/v1/videos"
KEY = os.environ.get("SUME_API_KEY", "")
assert KEY, "set SUME_API_KEY"
def call(url, body=None, idem=None):
headers = {"Authorization": f"Bearer {KEY}", "Content-Type": "application/json"}
if idem:
headers["Idempotency-Key"] = idem
data = json.dumps(body).encode() if body is not None else None
return urllib.request.urlopen(urllib.request.Request(url, data=data, headers=headers), timeout=60)
def make_clip(prompt, seconds=8, resolution="720p", idem="clip-001"):
body = {"model": "gemini-omni-flash-1.1", "prompt": prompt, "duration": seconds,
"resolution": resolution, "aspect_ratio": "16:9"}
job = json.load(call(BASE, body, idem))
while True:
time.sleep(30)
st = json.load(call(job["polling_url"]))
if st["status"] == "completed":
break
if st["status"] in ("failed", "cancelled"):
raise RuntimeError(st.get("error"))
with open("clip.mp4", "wb") as f:
f.write(call(st["unsigned_urls"][0]).read())
return st["usage"]["cost"]
print(make_clip("A paper boat crossing a puddle after rain, slow dolly in"))What changed from a Veo call
Google's deprecations page lists gemini-omni-1.1-flash as the replacement id. On Sume the id is gemini-omni-flash-1.1, bare, with no provider prefix. The other differences are in the limits of the model you are calling.
| Item | Before (Veo 3.1 preview) | After (Sume Omni row) |
|---|---|---|
| Model id | veo-3.1-generate-preview and two siblings | gemini-omni-flash-1.1 |
| Shutdown | October 22, 2026 (Google) | Read the live row at GET /v1/videos/models |
| Length | Chosen per Veo tier | 3 to 10 seconds, integer |
| Resolution | 720p, 1080p, 4K by tier | 360p, 720p, 1080p, 4K |
| Aspect ratio | Set in the Google request | 16:9 or 9:16 |
| Auth | Google credentials | Authorization: Bearer $SUME_API_KEY |
| Result | Google-hosted file | unsigned_urls[0], fetched with your key |
How the loop behaves
The submit returns immediately with a job id, a polling_url and the status pending. The job then moves to in_progress and ends in completed, failed or cancelled. Video generation usually takes between 30 seconds and several minutes according to the Sume docs, so a 30-second sleep means one to a handful of polls for a short clip. Only unsigned_urls[0] is read because this model returns one video per job.
To run it as a function in a larger program, import make_clip and pass a different idem for each clip. To run several clips at once, call it from threads or a job queue; each call is independent.
Things the script leaves out on purpose
It does not retry on a timeout, because a blind retry with a new key creates a second paid job. If the submit times out, call again with the same idem value and Sume returns the original job. It does not use a webhook; callback_url must be HTTPS, and a polling script is easier to run from a laptop.
Change the idem value for every new clip. Reusing a key returns the earlier job, which is the intended behavior for retries and a surprise for a new prompt.
Checking the price
The return value is the Sume billable amount for the job. For 8 seconds at 720p it should be $1.00, which is 8 x $0.125. If you ask for 1080p it should be $1.50. A different number means the clip came back at a different length, and the usage ledger at GET /v1/usage shows the reservation and capture.
Sources
Related posts
More in Developers
- Veo 3.1 previews end in 14 days: a dated checklist, Oct 8 to Oct 22
Google's three Veo 3.1 preview ids shut down on October 22, 2026. A day-by-day checklist from today, with the Sume model id and limits to test against.
- Vercel 800 s max duration: do you still need a Sume webhook?
Vercel Pro allows 800 s functions and a 30-minute beta. A Sume video job can still outlast one request, so use async or webhook mode and return in seconds.
- Vercel's 300 s default: does a 30 s Sume image sync call fit?
Yes: with fluid compute the default is 300 seconds, far above the 30-second Sume image wait. Branch on 200, 202 and 502, and poll video jobs instead.
- Verify a Sume Format webhook in Python: empty secret, stale timestamp
A Python check for the Sume format.run.terminal webhook: HMAC-SHA256 over timestamp.raw_body, a 5-minute window, and a refusal to run with an empty secret.
Written by Sume