Veo 3.1 returns one video per request: how to get 4 variants

Google's Veo 3.1 table says one video per request. To get four variants on Sume, send four requests with four Idempotency-Keys, and watch queue_full.

5 min readSume
All posts

Veo 3.1 gives you one video per request: Google's Veo model-features table lists "Videos per request" as 1 for Veo 3.1, Veo 3.1 Lite and Veo 3. To get four variants you send four requests. On Sume the same shape holds for its video models: one job per request, so four variants means four submits, each with its own Idempotency-Key.

Google's table is from its Gemini API: Generate videos with Veo, read 2026-10-02. The Sume side is from Video Router, Jobs and results and Generation admission. Sume does not list a Veo id, so the Sume example uses gemini-omni-flash-1.1, which the Video Router page documents.

Why a different Idempotency-Key for each variant?

Sume's docs say to reuse the same key only for the same operation and payload, and the Video generation page says a replay returns the original job. If all four requests shared one key you would get one job back four times. Give each variant its own key, and keep the key stable across retries of that variant.

for i in 1 2 3 4; do
  curl -s -X POST https://api.sume.com/v1/video-router/generate \
    -H "Authorization: Bearer $SUME_API_KEY" \
    -H "Content-Type: application/json" \
    -H "Idempotency-Key: lighthouse-variant-$i" \
    -d '{
      "model": "gemini-omni-flash-1.1",
      "prompt": "A lighthouse on a rocky cliff at dusk, waves crashing. Single continuous shot.",
      "resolution": "720p",
      "duration": 8,
      "aspect_ratio": "9:16",
      "mode": "async"
    }'
  echo
done

Will four identical prompts give four different videos?

Neither page promises that. Google's page says its seed only slightly improves determinism, and Sume's video models do not accept a seed. Vary the prompt a little per variant if you need guaranteed differences, such as the time of day or camera move, and compare the results.

What limits apply when you submit a batch?

Sume's admission page lists the failure modes: submit rate limits return 429 rate_limited, an unfunded estimate returns 402 insufficient_credits, and a full queue returns queue_full. For queue_full the docs say to stop adding work for that workspace, poll existing jobs until one is terminal, cancel queued jobs you no longer need, and retry with the same idempotency key after capacity opens.

Then poll each job with GET /v1/jobs/{id}/status until it is terminal, and read GET /v1/jobs/{id}/result for the completed ones.

Google's Veo guide and Sume's Generation admission page, read 2026-10-02.
ItemGoogle (Veo)Sume
Videos per request1One job per request
Submit throttleNot on this page429 rate_limited
Funds checkNot on this page402 insufficient_credits
Full queueNot on this pagequeue_full; stop adding, poll, retry with same key

How do I pick the winner?

Review all four results, keep the one you want, and note its job id. Variants you drop were still generated, so cap how many you run per brief.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume