Veo 3.1 returns one video per request: how to get 4 variants
Google's Veo 3.1 table says one video per request. To get four variants on Sume, send four requests with four Idempotency-Keys, and watch queue_full.

Veo 3.1 gives you one video per request: Google's Veo model-features table lists "Videos per request" as 1 for Veo 3.1, Veo 3.1 Lite and Veo 3. To get four variants you send four requests. On Sume the same shape holds for its video models: one job per request, so four variants means four submits, each with its own Idempotency-Key.
Google's table is from its Gemini API: Generate videos with Veo, read 2026-10-02. The Sume side is from Video Router, Jobs and results and Generation admission. Sume does not list a Veo id, so the Sume example uses gemini-omni-flash-1.1, which the Video Router page documents.
Why a different Idempotency-Key for each variant?
Sume's docs say to reuse the same key only for the same operation and payload, and the Video generation page says a replay returns the original job. If all four requests shared one key you would get one job back four times. Give each variant its own key, and keep the key stable across retries of that variant.
for i in 1 2 3 4; do
curl -s -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: lighthouse-variant-$i" \
-d '{
"model": "gemini-omni-flash-1.1",
"prompt": "A lighthouse on a rocky cliff at dusk, waves crashing. Single continuous shot.",
"resolution": "720p",
"duration": 8,
"aspect_ratio": "9:16",
"mode": "async"
}'
echo
doneWill four identical prompts give four different videos?
Neither page promises that. Google's page says its seed only slightly improves determinism, and Sume's video models do not accept a seed. Vary the prompt a little per variant if you need guaranteed differences, such as the time of day or camera move, and compare the results.
What limits apply when you submit a batch?
Sume's admission page lists the failure modes: submit rate limits return 429 rate_limited, an unfunded estimate returns 402 insufficient_credits, and a full queue returns queue_full. For queue_full the docs say to stop adding work for that workspace, poll existing jobs until one is terminal, cancel queued jobs you no longer need, and retry with the same idempotency key after capacity opens.
Then poll each job with GET /v1/jobs/{id}/status until it is terminal, and read GET /v1/jobs/{id}/result for the completed ones.
| Item | Google (Veo) | Sume |
|---|---|---|
| Videos per request | 1 | One job per request |
| Submit throttle | Not on this page | 429 rate_limited |
| Funds check | Not on this page | 402 insufficient_credits |
| Full queue | Not on this page | queue_full; stop adding, poll, retry with same key |
How do I pick the winner?
Review all four results, keep the one you want, and note its job id. Variants you drop were still generated, so cap how many you run per brief.
Sources
Related posts
More in Developers
- AI SDK stream cancel on disconnect: the Sume job keeps running
AI SDK 7.0.127 fixes stream cancellation when consumers disconnect. A cancelled stream does not cancel a Sume job: save the job id and read status later.
- AI SDK 7 tool search with deferred tools and Sume tools
AI SDK 7.0.127 lets a search() callback rank eligible deferred tools. Load a few Sume tools up front and fetch the rest by tools_schema on demand.
- Did the edit stay in its region? A pixel-diff check for GPT Image 2.5
After a GPT Image 2.5 edit, measure how much changed outside the area you meant to change. A Pillow script that diffs the result against the original.
- Verify a Sume avatar video webhook in Python (HMAC SHA-256)
A Python verifier for avatar video webhooks: timestamp tolerance, rotation-safe comparison, and a hard refusal when the signing secret is empty.
Written by Sume