Gemini Omni REST: output_video is SDK-only, read the steps array
Calling Gemini Omni over REST? interaction.output_video is SDK-only. Read the base64 video from the model_output step; on Sume you get a media URL instead.

When you call Gemini Omni Flash (gemini-omni-1.1-flash) with plain REST, interaction.output_video does not exist: Google's page says that convenience field is SDK-only, and that with raw REST you take the video from the steps array. Find the step whose type is model_output, then the content item whose type is video; its data is the base64-encoded MP4 and its mime_type is video/mp4.
This comes from Google's Gemini API: Generate and edit videos with Gemini Omni Flash, read 2026-10-02. Sume does not hand you base64: a Sume video job finishes with a media URL in its result, covered in Jobs and results and Video Router. Sume lists the model under its own id, gemini-omni-flash-1.1, so the two ids are not interchangeable.
Where is the video in a raw REST response?
Google shows the raw structure as an interaction object with an id, a status of completed, the model, and a steps array. The steps in its example are a user_input step, a thought step, and a model_output step. Only the last one carries the video.
Do not index the array by position. The example has three steps, but nothing on the page promises a fixed order or count, so loop over the steps and match on type. The snippet below does that against a response saved as a dict.
import base64
def video_bytes(interaction: dict) -> bytes:
for step in interaction["steps"]:
if step.get("type") != "model_output":
continue
for part in step["content"]:
if part.get("type") == "video":
return base64.b64decode(part["data"])
raise ValueError("no video part in any model_output step")
sample = {
"status": "completed",
"steps": [
{"type": "user_input", "content": [{"type": "text", "text": "..."}]},
{"type": "model_output", "content": [
{"type": "video", "mime_type": "video/mp4",
"data": base64.b64encode(b"fake-mp4").decode()}
]},
],
}
print(video_bytes(sample))What changes for videos larger than 4 MB?
Google's best-practices list says that for videos larger than 4 MB it recommends delivery="uri" in response_format to avoid payload size limits. Inline base64 is therefore a choice for small outputs, not a rule for every clip.
Google also says that setting store=false means the generated video cannot be edited in later turns with previous_interaction_id. If you plan to edit, keep store on and keep the interaction id.
What does the same step look like on Sume?
On Sume you submit to POST /v1/video-router/generate with mode: "async", then poll the job and read GET /v1/jobs/{id}/result. The docs describe completed jobs as carrying result.artifacts[], each with a url, a type, and a content_type, and say to use the Sume media URLs from the result because raw provider URLs are not public API outputs.
That removes the decode step: you download or hand on a URL. It also means a base64 parser written for Google's REST shape will not work on a Sume result, and the reverse.
| Question | Gemini API (REST) | Sume |
|---|---|---|
| Model id | gemini-omni-1.1-flash | gemini-omni-flash-1.1 |
| Where the video is | steps[] item model_output, content type video | result.artifacts[].url after the job completes |
| Form of the video | Base64 data, or a URI with delivery="uri" | A Sume media URL |
output_video shortcut | SDK-only | Not part of the job result |
What should I check before shipping?
Three habits keep a REST client honest.
- Match on
type, not on array position. - Treat a
statusother thancompletedas no video. - On Sume, wait for
terminalandresult_readybefore reading the result;GET /v1/jobs/{id}/resultanswers409 job_not_completeduntil then.
Sources
Related posts
More in Models
- Gemini Omni [# Sources] and [# References] tags vs Sume fields
Gemini Omni binds media to roles with tags like <FIRST_FRAME> and [# Sources ...]. Sume uses request fields instead: image_url, end_image_url, reference lists.
- Gemini Omni and Veo in Korean: only English is fully supported
Google's Omni and Veo pages say English is fully supported; other languages aren't evaluated. For Korean prompts, describe in English and quote on-screen text.
- Gemini Omni can't use a YouTube link as its source video
Google lists YouTube videos as an unsupported media source for Gemini Omni Flash. On Sume, video_url is a URI field; use a direct link to the clip file.
- GPT Image 2.5 mask_url edit: change one region, keep the rest
GPT Image 2.5 on Sume accepts a public mask_url alongside input_references. Build a mask with Pillow, host it, and edit only one region of a photo.
Written by Sume