Gemini Interactions API image output vs Sume /v1/images
Gemini's Interactions API returns the image on interaction.output_image. Sume's /v1/images returns data[].url, or a 202 job envelope. Check the status code.

With the Gemini API's Interactions API you read the generated image from the interaction.output_image property. On Sume you call POST /v1/images and read data[].url, a Sume-hosted signed URL; if the render outlasts 30 seconds the call returns 202 and a job envelope instead.
Gemini facts are from Google's image generation page; Sume facts from the Image API docs, read 2026-10-01.
How does Gemini return the image?
Google's page shows client.interactions.create with model gemini-3.1-flash-image and says you retrieve generated image data with the interaction.output_image property, which returns the last generated image block. The Python sample base64-decodes output_image.data and writes it to a file.
How does Sume return it?
The result payload is data[].url (Sume-hosted, signed) rather than inline base64, which keeps responses small. The Google image family is in the catalog as google/nano-banana-2, and the bare router id nano-banana-2 is accepted as an alias.
| Question | Gemini Interactions API | Sume /v1/images |
|---|---|---|
| Where is the image? | interaction.output_image | data[].url |
| Encoding | Base64 data in the sample | Signed URL |
| Slow render | Not covered on the page | 202 with a job envelope |
What if the response is 202?
POST /v1/images blocks for up to 30 seconds and returns 200 with the images. If generation is still running, or you send mode: "async", you get 202 with status_url and result_url. The docs are explicit: check the status code, not the body shape. Then poll GET /v1/jobs/{id}/status and fetch GET /v1/jobs/{id}/result; see Jobs and results.
const res = await fetch("https://api.sume.com/v1/images", {
method: "POST",
headers: {
Authorization: "Bearer " + process.env.SUME_API_KEY,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "google/nano-banana-2",
prompt: "a nano banana dish in a fancy restaurant",
}),
});
const body = await res.json();
if (res.status === 200) console.log(body.data[0].url);
else if (res.status === 202) console.log(body.data.result_url);Which model id do I send?
Use the Sume catalog id, not the Gemini id. Sume serves every catalog model through a single sume endpoint in v1. For a similar comparison on the video side, see previous interaction id vs a job per request.
Sources
Related posts
More in Developers
- Gemini last_event_id stream resume vs Sume: no SSE, poll status_url
Gemini background interactions can resume a dropped stream from the last event. Sume has no stream to resume: poll status_url and read events_url snapshots.
- Omni 1.1 Flash start and end frame API on Sume
Omni 1.1 Flash can render between two keyframes, including looping clips. On Sume, Omni takes image_url plus end_image_url; frame_images covers other models.
- Gemini Omni videos over 4MB: poll to ACTIVE vs a Sume result URL
Gemini Omni returns large videos as a URI you poll until ACTIVE. Sume results are durable media.sume.com URLs with no ACTIVE state and no expiry to manage.
- Gemini Omni Flash video edit: 5 reference images vs Sume
Runway's Gemini Omni Flash video-to-video edit accepts up to 5 reference images. On Sume, a video_url edit cannot be combined with any reference field.
Written by Sume