Nano Banana video to image: poster from a video via Sume stills
Google's Gemini API takes a video as context for a thumbnail or poster. Sume's images API takes image references only, so grab a still frame first.
Google's Gemini API can take a video, including a public YouTube URL, as context for a new thumbnail or poster. Sume's images endpoint does not document a video input: input_references takes image URLs. To make a poster from a video on Sume, export a still frame, host it at a public HTTPS URL and pass it as a reference.
Google's feature is described on its image generation page; Sume's inputs on Image generation, both read 2026-10-01.
What does Google's video-to-image feature do?
Google's page says video-to-image generation lets you create new images using a video's context as a multimodal reference, for video thumbnails, cinematic posters, summary infographics or new artwork inspired by a scene. The model analyzes frames in context and combines them with your text prompt. You can pass public YouTube URLs in the request or upload local video files. The page limits video inputs to Gemini 3.1 Flash Image and Gemini 3.1 Flash Lite Image.
What does Sume accept as input?
The input_references parameter is listed as an array of reference images for image-to-image. Reference URLs must be public HTTPS; localhost, private-network and non-HTTPS URLs are rejected before submission. Whether a given model takes references at all depends on its catalog descriptor, and the catalog includes google/nano-banana-2.
| Question | Google Gemini API | Sume POST /v1/images |
|---|---|---|
| Video as context | Yes, on the two 3.1 image models | Not documented |
| YouTube URL | Public URLs accepted | Not documented |
| Image input | Yes | input_references, public HTTPS |
How do I get a poster from a video on Sume?
Pick the frame that best carries the scene, export it as a PNG or JPEG, and host it at a public HTTPS URL. Send it in input_references with a prompt that says what the poster should add, such as title space or a different crop. On edit and image-to-image calls the docs advise aspect_ratio: "auto" to match the reference; omitting the field is not the same. See also how references behave on Nano Banana.
What do I lose by using a single frame?
Google's model reads several frames and the events between them; a single still carries one moment. If the poster should combine two scenes, export two frames and pass both as references, within the per-model limit shown in the catalog. Sume does not describe any frame extraction of its own on this endpoint, so that step stays in your pipeline.
Sources
Related posts
More in Models
- Runway video_to_hdr alpha and ProRes 4444: what Sume returns
Runway's /v1/video_to_hdr now keeps source alpha and delivers ProRes 4444. Sume docs describe no alpha or ProRes output; a completed video job is a download.
- Speechmatics Linden for voice agents vs Sume batch STT jobs
Speechmatics announced Linden for voice agents. Sume STT is a bounded batch job with a webhook, not a live stream, so live agents need streaming.
- Synthesia Interactive Avatar API vs Sume rendered avatar clips
Synthesia headlines a live Interactive Avatar API. Sume avatar video is script-driven and rendered as a job: submit, poll, then fetch the clip.
- An OpenRouter-compatible video API: sume/auto or a pinned model
Sume's POST /v1/videos follows OpenRouter's video generation API field for field. Let sume/auto pick the model, or pin a catalog id like seedance-2.5.
Written by Sume