Cloud Run 504 after 300 seconds: return 202 for Sume video jobs

Cloud Run ends a request after 5 minutes by default with a 504. Don't wait for a render in the request: submit to Sume, return 202, then poll or use a webhook.

5 min readSume
All posts

Cloud Run has a default request timeout of 5 minutes (300 seconds) and a maximum of 60 minutes, and it returns a 504 when your handler is still running at the limit. A video render can take longer than that, so your endpoint should submit the job to Sume, return 202 with the job id at once, and learn the result later by polling or a webhook.

The numbers

Google's docs say the connection closes and an error 504 is returned when no response is sent in the configured time. You set it in the console (1 to 3600 seconds), with gcloud run services update SERVICE --timeout=TIMEOUT, or with timeoutSeconds in YAML. For timeouts above 15 minutes Google recommends retries and graceful reconnects. Sume's Videos guide says generation is asynchronous and can take several minutes depending on model, resolution and load.

Cloud Run request timeouts (read 2026-10-05)
SettingValue
Default timeout5 minutes (300 seconds)
Maximum timeout60 minutes (3600 seconds)
Status on timeout504
Long timeoutsAbove 15 minutes, add retries and reconnect handling

Why raising the timeout is the wrong fix

A 60-minute timeout keeps a container and a client socket busy for the whole render, and a dropped connection loses the job id. If the submit used an Idempotency-Key, a retry is safe, but you still pay for an instance sitting idle. The cheaper pattern is two short requests: one to start, one to deliver.

A handler that returns in under a second

This stdlib server starts a job, replies 202 with the id and polling URL, and exits the request. It uses the key from the request so a retry cannot create a second job.

import os, json, requests
from http.server import BaseHTTPRequestHandler, HTTPServer

class H(BaseHTTPRequestHandler):
    def do_POST(self):
        n = int(self.headers.get("content-length", 0))
        body = json.loads(self.rfile.read(n) or b"{}")
        r = requests.post(
            "https://api.sume.com/v1/videos", timeout=30,
            json={"model": "gemini-omni-flash-1.1", "prompt": body["prompt"],
                  "callback_url": os.environ["SUME_CALLBACK_URL"]},
            headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
                     "Idempotency-Key": body["request_key"]})
        out = json.dumps(r.json()).encode()
        self.send_response(202 if r.ok else r.status_code)
        self.send_header("content-type", "application/json")
        self.end_headers()
        self.wfile.write(out)

HTTPServer(("", int(os.environ.get("PORT", 8080))), H).serve_forever()

Checklist for the two-service pattern

Keep the timeout at its default unless a handler truly needs longer, and treat any handler that waits on a render as a design bug.

  • Service one starts jobs and returns in under a second.
  • Service two receives the signed webhook and replies with a 2xx before doing slow work.
  • Both services store the job id and ignore repeats.
  • A scheduled poll closes the gap for events that never arrive.

Delivering the result

callback_url must be HTTPS and receives a signed job event when the job reaches a terminal state, with x-sume-webhook-signature headers. Sume makes up to 10 delivery attempts, 30 seconds apart by default, with a 10-second timeout per attempt, so the receiving service must reply quickly and do its slow work afterwards. Keep polling as a fallback, because delivery is an optimization and not the only recovery path. See Webhooks.

Sources

Related posts

More in Integrations

All Integrations posts

Written by Sume