Omni takes JPEG and PNG: gate each reference image URL in Python

Google lists JPEG and PNG for Gemini Omni image input. Check each reference_image_urls entry with a HEAD request, then submit the clean list through Sume.

4 min readSume
All posts

Send only JPEG and PNG files as Gemini Omni reference images, and check the content type of each URL before you pay for a job. Google's Omni page lists those two image types, and Sume's Video Router takes up to ten URLs in reference_image_urls, so one bad entry in a ten-image list can waste the whole request.

The Sume docs do not publish a format list for Omni references, so the safe rule is the stricter one: the vendor's.

What does each side say?

Note that the model id differs between the two sites. Google writes gemini-omni-1.1-flash, and Sume's catalog id is gemini-omni-flash-1.1. Use the Sume spelling in Sume requests.

Omni reference inputs, vendor page vs Sume docs (read 2026-10-09)
ItemGoogle's Omni pageSume Video Router
Model idgemini-omni-1.1-flashgemini-omni-flash-1.1
Image typesJPEG, PNGNot listed in the docs
Reference imagesImages as inputreference_image_urls, up to 10
Prompt tagsNot covered<IMAGE_REF_0>, 0-based, in list order

How do I filter the list?

A HEAD request is cheap, but some hosts answer 405 to it. Treat a failed HEAD as a skip and log the URL, rather than guessing. The sample then submits through POST /v1/video-router/generate with an Idempotency-Key, so a retried script does not bill twice.

import os, requests

OK = {"image/jpeg", "image/png"}

def usable(urls):
    good = []
    for u in urls[:10]:
        h = requests.head(u, allow_redirects=True, timeout=10)
        kind = h.headers.get("content-type", "").split(";")[0].strip().lower()
        if h.ok and kind in OK:
            good.append(u)
        else:
            print("skip", u, h.status_code, kind)
    return good

refs = usable(os.environ["REFS"].split(","))
if not refs:
    raise SystemExit("no usable reference images")
r = requests.post(
    "https://api.sume.com/v1/video-router/generate",
    headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
             "Idempotency-Key": os.environ["KEY"]},
    json={"model": "gemini-omni-flash-1.1", "mode": "async",
          "prompt": "<IMAGE_REF_0> on a desk, slow push-in",
          "reference_image_urls": refs, "resolution": "720p",
          "duration": 6, "aspect_ratio": "16:9"},
    timeout=30)
print(r.status_code, r.json()["data"]["request_id"])

What if the prompt names a tag that was dropped?

Tags count by position in the list you send. If the filter removes the first image, <IMAGE_REF_0> now points at the old second image. Rebuild the prompt from the filtered list, or fail the script when any entry is skipped, so the picture a tag refers to never changes silently.

Why not convert WebP in the same script?

You can, but Sume needs a public HTTPS URL, so a converted file must be re-hosted before you submit. Keep the gate and the conversion as separate steps, and read each reference image's Content-Type again after you upload it.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume