Gemini API webhooks for batch jobs: keep a poll backup, as on Sume
Gemini API release notes say webhooks replace polling for Batch and long-running operations. What changes in a client, and why Sume keeps a polling backup.

The Gemini API changelog says webhook support replaces polling for the Batch API and long-running operations, so a client can wait for a push instead of looping on a status call. The sound way to adopt it is the one Sume documents for its own job webhooks: take the push, store it first, and keep a poll available for the events that never arrive.
This post compares the two ideas at the level of client design. It does not describe Gemini's webhook format, which is on Google's pages.
What the changelog lists
Only the webhook entry is used.
| Item | What the page says |
|---|---|
| Webhooks | Replace polling for the Batch API and long-running operations |
What changes in a client
A polling client owns the loop: it submits, then asks until the answer changes. A webhook client owns a public endpoint: it submits, stores the operation id, and waits to be called. The second client needs a stable HTTPS URL, a way to verify who is calling, and a way to handle the same call twice.
Sume's job webhooks make the same trade. A submit with mode: "webhook" and a public HTTPS webhook_url returns 202 with the job envelope and polling URLs, and the callback is stored. Localhost, private-network and non-HTTPS URLs are rejected.
What Sume does about delivery
The delivery rules are explicit, which makes them a useful template for any webhook client.
| Property | Value |
|---|---|
| Events | job.completed, job.failed, job.canceled; terminal only |
| Signature | HMAC SHA 256 over <timestamp>.<raw_body>, header sume-v1=<hex> |
| Retries | Up to 10 attempts in total |
| Spacing | Fixed delay, 30 seconds by default |
| Timeout | 10 seconds per attempt |
| Receiver key | Use job_id as the idempotency key |
Why keep the poll
Ten refused attempts leave a failed delivery and a job that still reached its real terminal state. Sume's docs call delivery an optimization and never the only recovery path, and they tell you to keep status_url polling for the events that never arrive. A client that deletes its poll loop the day it adopts webhooks has no recovery when the endpoint is down.
The poll can be slow and cheap. Check once on a long interval for jobs that have no stored event after a deadline, and use next_poll_after_seconds from the status response when it is present.
A receiver that survives repeats
The same four steps apply to any provider's webhook.
- Verify the signature on the raw body before parsing it. Reject callbacks whose timestamp falls outside a replay window; Sume's docs call five minutes a reasonable default.
- Store the event keyed by the job or operation id. A repeat is then a no-op.
- Answer
2xxafter storing, not after the slow work. - Let a separate worker, and a slow poll, finish what the webhook missed.
Run webhooks are a different surface
Sume has a second webhook family for Format, Action and Agent Completion runs, with events format.run.terminal, action.run.terminal and agent.run.terminal. The signature scheme is identical, so one verifier covers both, but the payload is a full run receipt, not a job result. Pick the handler by event name.
The same applies when you mix providers: route by source and event first, then verify, then store.
Sources
Related posts
More in Comparisons
- Gemini Omni Flash 1,240 vs Seedance 2.0 1,225: what 15 Elo means
Hedra lists Gemini Omni Flash first at 1,240 Elo and Seedance 2.0 4K second at 1,225. A 15-point gap is about a 52% win rate; check limits before choosing.
- Google's video default is Omni Flash: three cases where Veo 3.1 fits
Google's video docs call Gemini Omni Flash the default and keep Veo 3.1 for extension, frame-specific and image-directed jobs. What each maps to on Sume.
- Google Ads built-in image and video generation vs a Sume pipeline
Demand Gen can generate up to 20 images per prompt and auto-build video from a logo, two images and two texts. Where a separate Sume pipeline differs.
- Image reference limits: 10 on ElevenLabs, 16 on Sume, 5 Ideogram 4.5
ElevenLabs lists 10 references for GPT Image 2.5. Sume's docs list 16 for the same models and 5 total for Ideogram 4.5. A table, a count guard and an edit rule.
Written by Sume