One Gemini billing cap pauses every linked project: guard the batch

Google pauses all projects on a billing account at the tier cap. Add a balance check and idempotent submits so one Omni batch cannot stall the rest.

5 min readSume
All posts

On the Gemini API, a spend cap is applied per billing account: Google's billing page says that after the cumulative total reaches the tier limit, service is paused for all projects linked to that billing account until the next billing cycle. A runaway Omni video batch in a test project can therefore stall your production text calls. Guard it with your own counter, one job per idempotency key, and a hard stop on payment errors.

The failure mode in Google's words

The billing page separates account-level monthly caps from project-level caps, and warns that project caps can be overshot for around a 10 minute latency period. For prepaid credits it says that when the Prepay balance hits $0, all API keys in all projects linked to that billing account stop working simultaneously. Omni Flash is billed at about $0.10 per second at 720p on the pricing page, so a retry loop that regenerates ten-second clips burns about $1.00 a pass.

The common mistake is a retry on a slow job that submits a second paid generation. If your worker times out and resubmits a paid generation, each submit is a separate paid request unless you deduplicate on your side.

Guard rails, side by side

Spend guard rails for a video batch (Google billing page and Sume docs, read 2026-10-04)
ControlGemini APISume
Scope of the stopWhole billing account, all linked projectsWorkspace balance, per submit
Failure signalService pauses until next cycle402 insufficient_credits at submit
OvershootProject caps lag about 10 minutesReservation made before provider work
Duplicate protectionYour codeIdempotency-Key header on paid submits

A client-side stop for Sume jobs

This sketch checks the balance before each paid submit and refuses to continue when it falls below one clip. It reads the balance from GET /v1/balance, sends a stable Idempotency-Key per item and returns the job id. Sume documents that paid submits can be retried with the same key, and that a 409 idempotency_conflict means the key was reused for a different payload (Generation admission).

Keep the key derived from your own item id, not a random value, so a crashed worker that restarts submits the same request again instead of a new one.

import os
import requests

API = "https://api.sume.com"
HEADERS = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
CLIP_USD = 1.25  # 10 s at 720p, list x 1.25

def balance_usd() -> float:
    r = requests.get(f"{API}/v1/balance", headers=HEADERS, timeout=30)
    r.raise_for_status()
    micros = r.json()["data"]["balance"]["available_amount_usd_micros"]
    return micros / 1_000_000

def submit(item_id: str, prompt: str) -> str:
    if balance_usd() < CLIP_USD:
        raise SystemExit("balance below one clip; stopping the batch")
    r = requests.post(
        f"{API}/v1/video-router/generate",
        headers={**HEADERS, "Idempotency-Key": f"omni-batch-{item_id}"},
        json={"model": "gemini-omni-flash-1.1", "prompt": prompt,
              "resolution": "720p", "duration": 10,
              "aspect_ratio": "9:16", "mode": "async"},
        timeout=30,
    )
    r.raise_for_status()
    return r.json()["data"]["job"]["id"]

What this does not cover

A balance check is a snapshot; other jobs may reserve funds between your check and your submit, and a 402 can still arrive, so handle it as a stop, not a retry (Errors and credits). It also does not cap spend on Google; if you use both routes, keep separate billing accounts per workload so one cap cannot pause the other.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume