Gemini Omni REST: output_video is SDK-only, read the steps array

Calling Gemini Omni over REST? interaction.output_video is SDK-only. Read the base64 video from the model_output step; on Sume you get a media URL instead.

5 min readSume
All posts

When you call Gemini Omni Flash (gemini-omni-1.1-flash) with plain REST, interaction.output_video does not exist: Google's page says that convenience field is SDK-only, and that with raw REST you take the video from the steps array. Find the step whose type is model_output, then the content item whose type is video; its data is the base64-encoded MP4 and its mime_type is video/mp4.

This comes from Google's Gemini API: Generate and edit videos with Gemini Omni Flash, read 2026-10-02. Sume does not hand you base64: a Sume video job finishes with a media URL in its result, covered in Jobs and results and Video Router. Sume lists the model under its own id, gemini-omni-flash-1.1, so the two ids are not interchangeable.

Where is the video in a raw REST response?

Google shows the raw structure as an interaction object with an id, a status of completed, the model, and a steps array. The steps in its example are a user_input step, a thought step, and a model_output step. Only the last one carries the video.

Do not index the array by position. The example has three steps, but nothing on the page promises a fixed order or count, so loop over the steps and match on type. The snippet below does that against a response saved as a dict.

import base64


def video_bytes(interaction: dict) -> bytes:
    for step in interaction["steps"]:
        if step.get("type") != "model_output":
            continue
        for part in step["content"]:
            if part.get("type") == "video":
                return base64.b64decode(part["data"])
    raise ValueError("no video part in any model_output step")


sample = {
    "status": "completed",
    "steps": [
        {"type": "user_input", "content": [{"type": "text", "text": "..."}]},
        {"type": "model_output", "content": [
            {"type": "video", "mime_type": "video/mp4",
             "data": base64.b64encode(b"fake-mp4").decode()}
        ]},
    ],
}
print(video_bytes(sample))

What changes for videos larger than 4 MB?

Google's best-practices list says that for videos larger than 4 MB it recommends delivery="uri" in response_format to avoid payload size limits. Inline base64 is therefore a choice for small outputs, not a rule for every clip.

Google also says that setting store=false means the generated video cannot be edited in later turns with previous_interaction_id. If you plan to edit, keep store on and keep the interaction id.

What does the same step look like on Sume?

On Sume you submit to POST /v1/video-router/generate with mode: "async", then poll the job and read GET /v1/jobs/{id}/result. The docs describe completed jobs as carrying result.artifacts[], each with a url, a type, and a content_type, and say to use the Sume media URLs from the result because raw provider URLs are not public API outputs.

That removes the decode step: you download or hand on a URL. It also means a base64 parser written for Google's REST shape will not work on a Sume result, and the reverse.

Google's Omni guide and Sume's Jobs and results page, read 2026-10-02.
QuestionGemini API (REST)Sume
Model idgemini-omni-1.1-flashgemini-omni-flash-1.1
Where the video issteps[] item model_output, content type videoresult.artifacts[].url after the job completes
Form of the videoBase64 data, or a URI with delivery="uri"A Sume media URL
output_video shortcutSDK-onlyNot part of the job result

What should I check before shipping?

Three habits keep a REST client honest.

  • Match on type, not on array position.
  • Treat a status other than completed as no video.
  • On Sume, wait for terminal and result_ready before reading the result; GET /v1/jobs/{id}/result answers 409 job_not_completed until then.

Sources

Related posts

More in Models

All Models posts

Written by Sume