Cloud Tasks settings for Sume video jobs: pace submits, retry 429

Cloud Tasks defaults to 500 dispatches a second, far above a Pro plan's 24 accepted jobs. Pace submits and retry 429 queue_full with the same Idempotency-Key.

5 min readSume
All posts

Cloud Tasks example defaults are 500 dispatches per second and 1,000 concurrent dispatches, while a Sume Pro workspace accepts 24 paid generation jobs at a time (4 processing and 20 queued). Set a low dispatch rate on the queue, and make the task handler retry a 429 queue_full with the same Idempotency-Key.

What a task gates

A task that submits to Sume finishes in milliseconds, because the API returns 202 and the render runs on its own. So maxConcurrentDispatches does not limit renders; it limits submit calls. What limits renders is Sume's admission rule: jobs wait as queued until a processing slot opens, and queue_full appears only when both processing and queue capacity are used.

Cloud Tasks queue knobs and Sume limits (read 2026-10-05)
Knob or limitValueEffect
maxDispatchesPerSecond500 in Google's exampleSet to 1 or 2 for paid video submits
maxConcurrentDispatches1,000 in Google's exampleCaps parallel submit calls only
maxAttempts / minBackoff100 default (-1 is unlimited) / 0.100s defaultUse finite attempts and 30s+ backoff
Sume Pro4 processing, 20 queued, 24 accepted429 queue_full when full

The queue

Google's docs describe exponential backoff: the retry interval starts at the minimum, doubles maxDoublings times, then grows linearly up to the maximum. Both maxAttempts and maxRetryDuration must be reached to stop retries. Pick numbers that outlast a typical render so a queued workspace drains.

import subprocess

cmd = ["gcloud", "tasks", "queues", "create", "sume-video-submits",
       "--max-dispatches-per-second=1",
       "--max-concurrent-dispatches=2",
       "--max-attempts=20",
       "--min-backoff=30s",
       "--max-backoff=600s"]
print(" ".join(cmd))
# Uncomment after gcloud auth to create the queue:
# subprocess.run(cmd, check=True)

The handler's answer table

A task is retried when your handler returns a non-2xx. So map Sume's answers: 202 becomes your 200 (done); 429 queue_full or rate_limited becomes a 429 or 503 for retry (honour retry-after when present); 402 insufficient_credits and 400 must become a 2xx with an alert, because retries cannot fix them. Always send the same Idempotency-Key per task, because Sume treats reuse for an exact retry as safe and a different body as a 409 conflict.

Failure modes to test

Run these four against a test queue with a tiny balance and a low plan before the real batch, and read the status and code of each answer.

  • A 402 for low balance: the task should stop and alert, not retry.
  • A 400 for a bad body: the task should stop and log the payload.
  • A 429 queue_full: the task should retry later with the same key.
  • A network error before the response: the task retries, and the key prevents a second job.

Sizing a batch

Pro accepts 24 jobs at a time. A batch of 100 clips released at 1 per second will hit the cap in about 24 seconds, then see queue_full until jobs finish. That is fine if the retries are cheap, but a rate near your finish rate wastes fewer calls. Read generation_limits in the submit response, and the Generation admission page, for the live values.

Sources

Related posts

More in Integrations

All Integrations posts

Written by Sume