Nano Banana 2.1 PDF and video grounding: Sume takes image URLs

fal's Nano Banana 2.1 Edit page lists video and PDF grounding. Sume's Image API takes public HTTPS image URLs only, so render pages first.

5 min readSume
All posts

Sume's Image API cannot take a PDF or a video as grounding for Nano Banana 2.1. Its input_references entries are image URLs: each needs image_url.url, it must be public HTTPS, and Sume rejects localhost, private-network and non-HTTPS addresses before submission. The fal page for Nano Banana 2.1 Edit (read 2026-10-08) lists "optional video/PDF grounding for context-based generation" among the model's editing features, and Sume's docs list no field for it.

The workable path is to turn the pages you care about into images yourself and pass those. The row accepts up to 10 references on Sume, against the 14 the fal edit page lists, so a long document needs a selection step first.

What each side lists

The comparison is limited to input types and counts. Both pages were read on 2026-10-08, and Sume's side comes from the Image API docs and the request validator on origin/main.

Nano Banana 2.1 grounding and reference inputs, fal vs Sume (read 2026-10-08)
Inputfal Nano Banana 2.1 Edit pageSume Image API
Reference imagesUp to 14 in one requestUp to 10 in input_references
PDF groundingListed as optionalNo field; render pages to images
Video groundingListed as optionalNo field; export frames to images
Mask editingSemantic masking and mask-based targeting listedmask_url is accepted by ChatGPT Image 2.5 only
Characters kept consistentUp to 4 people listedSent through references; Sume adds no tracking field

Render pages, then send them

Export the pages as PNG or JPEG, upload them somewhere with a public HTTPS URL, and put each in an input_references entry in reading order. Say in the prompt what each page is for, because the model sees images, not page numbers. Sume answers 200 with the image when the work finishes inside 30 seconds and 202 with a job envelope when it does not; 2K and 4K edits are the likeliest to take the second path.

import os, requests

body = {
    "model": "google/nano-banana-2.1",
    "prompt": "Redraw page 1 as a clean 16:9 slide, keep the chart colors from page 2",
    "input_references": [
        {"type": "image_url", "image_url": {"url": "https://example.com/p1.png"}},
        {"type": "image_url", "image_url": {"url": "https://example.com/p2.png"}},
    ],
    "aspect_ratio": "16:9",
    "resolution": "2K",
}
r = requests.post(
    "https://api.sume.com/v1/images",
    json=body,
    headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
    timeout=60,
)
if r.status_code == 202:
    print("still running:", r.json()["data"]["status_url"])
else:
    r.raise_for_status()
    print(r.json()["data"][0]["url"], r.json()["usage"]["cost"])

Choosing which pages to send

Ten slots go quickly. A cover, one data page and one style page usually carry more than ten dense pages. If the PDF is text, extract the text and put the facts in the prompt instead of the image slots. If the source is a video, take a few key frames, not every second.

Cost stays per image. Sume bills a flat card per resolution tier (see the linked posts), and cost_usd × n is the charge for the call. Extra references do not add input tokens on this row.

One practical rule keeps the pipeline honest: whatever you send as a reference should be something you would be willing to look at yourself. If a rendered page is unreadable at the size you upload it, the model will not read it either, so export at a size where the text and chart labels are legible, and keep each file within what your host serves over plain HTTPS.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume