Gemini batch inline 20 MB or 2 GB file vs Sume's 100 inline items

Gemini Batch takes inline requests under 20 MB or a file up to 2 GB. Sume bulk takes 1 to 100 inline items. How to size a holiday video batch for each.

5 min readSume
All posts

How big can one batch be on Gemini compared with Sume? The Gemini Batch API guide (read 2026-10-04) says inline requests are for total request size under 20 MB, while an input file can be up to 2 GB. Sume has only the inline form: a bulk queue takes items, an array of 1 to 100 run bodies, in one request.

If your holiday job is 10,000 SKUs, that is 100 Sume queues, not one. If it is 80, it is one queue and it never needs a file.

What counts toward size on the Sume side

Each bulk item is the same body as a single run. The Create run page caps a run request at 4 MiB (413 payload_too_large), and caps input at 64 top-level keys and 2 MiB. The bulk page does not publish a separate cap for the whole envelope, so check your serialized body yourself instead of assuming 100 large items fit.

Send media by URL, not by value. The docs say an attachment's image_url is fetched when the run is created, so it must be reachable without auth. A broken URL fails the create with a 4xx or 5xx before any queue exists.

Sizing the plan

Batch size rules, Gemini page and Sume docs read 2026-10-04
Gemini Batch APISume bulk run
Inline formTotal request size under 20 MBThe only form: 1 to 100 items
File formInput file up to 2 GBNot offered
Per-run bodySet by your requests4 MiB per run request, input 2 MiB
ParallelismService decidesYou set concurrency 1 to 16

Chunk and key each queue

Chunk at 100, give every chunk its own Idempotency-Key built from the campaign and the chunk number, and keep the returned frq_ ids in a table. Replaying a spent key returns the old queue, which makes a crashed uploader safe to restart. The code below does the chunking and prints the key per chunk.

Because there is no list-queues endpoint, losing the queue ids means losing the way to read them. Persist them as soon as each 202 returns.

def chunks(rows, size=100):
    for i in range(0, len(rows), size):
        yield i // size, rows[i : i + size]


skus = [f"sku-{n:04d}" for n in range(250)]
campaign = "holiday-2026"
plan = []
for n, part in chunks(skus):
    key = f"{campaign}-chunk-{n:03d}"
    plan.append((key, len(part)))
assert sum(c for _, c in plan) == len(skus)
for key, count in plan:
    print(key, count)

When a file would have helped, and what to do instead

A 2 GB input file exists because some batch jobs carry large prompts or embedded media. A Sume item is deliberately small: media goes by HTTPS URL, and the docs say Sume fetches attachments at create time, checks their real type and size, and copies them into durable storage. A private or hotlink-protected URL fails the create with attachment_fetch_failed before a queue exists, so test a sample URL first.

The practical plan is a loop: split rows into chunks of 100, validate every media URL with a HEAD or GET from your own server, post each chunk with its own key, and store the queue ids. If one chunk fails validation, only that chunk waits; the others keep running.

That structure also fits the rate limits: creating a queue spends the write budget, while polling spends the separate, larger read budget, so a 250-SKU job is three writes and as many reads as your backoff allows.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume