Cloud Run 504 after 300 seconds: return 202 for Sume video jobs
Cloud Run ends a request after 5 minutes by default with a 504. Don't wait for a render in the request: submit to Sume, return 202, then poll or use a webhook.

Cloud Run has a default request timeout of 5 minutes (300 seconds) and a maximum of 60 minutes, and it returns a 504 when your handler is still running at the limit. A video render can take longer than that, so your endpoint should submit the job to Sume, return 202 with the job id at once, and learn the result later by polling or a webhook.
The numbers
Google's docs say the connection closes and an error 504 is returned when no response is sent in the configured time. You set it in the console (1 to 3600 seconds), with gcloud run services update SERVICE --timeout=TIMEOUT, or with timeoutSeconds in YAML. For timeouts above 15 minutes Google recommends retries and graceful reconnects. Sume's Videos guide says generation is asynchronous and can take several minutes depending on model, resolution and load.
| Setting | Value |
|---|---|
| Default timeout | 5 minutes (300 seconds) |
| Maximum timeout | 60 minutes (3600 seconds) |
| Status on timeout | 504 |
| Long timeouts | Above 15 minutes, add retries and reconnect handling |
Why raising the timeout is the wrong fix
A 60-minute timeout keeps a container and a client socket busy for the whole render, and a dropped connection loses the job id. If the submit used an Idempotency-Key, a retry is safe, but you still pay for an instance sitting idle. The cheaper pattern is two short requests: one to start, one to deliver.
A handler that returns in under a second
This stdlib server starts a job, replies 202 with the id and polling URL, and exits the request. It uses the key from the request so a retry cannot create a second job.
import os, json, requests
from http.server import BaseHTTPRequestHandler, HTTPServer
class H(BaseHTTPRequestHandler):
def do_POST(self):
n = int(self.headers.get("content-length", 0))
body = json.loads(self.rfile.read(n) or b"{}")
r = requests.post(
"https://api.sume.com/v1/videos", timeout=30,
json={"model": "gemini-omni-flash-1.1", "prompt": body["prompt"],
"callback_url": os.environ["SUME_CALLBACK_URL"]},
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": body["request_key"]})
out = json.dumps(r.json()).encode()
self.send_response(202 if r.ok else r.status_code)
self.send_header("content-type", "application/json")
self.end_headers()
self.wfile.write(out)
HTTPServer(("", int(os.environ.get("PORT", 8080))), H).serve_forever()
Checklist for the two-service pattern
Keep the timeout at its default unless a handler truly needs longer, and treat any handler that waits on a render as a design bug.
- Service one starts jobs and returns in under a second.
- Service two receives the signed webhook and replies with a 2xx before doing slow work.
- Both services store the job id and ignore repeats.
- A scheduled poll closes the gap for events that never arrive.
Delivering the result
callback_url must be HTTPS and receives a signed job event when the job reaches a terminal state, with x-sume-webhook-signature headers. Sume makes up to 10 delivery attempts, 30 seconds apart by default, with a 10-second timeout per attempt, so the receiving service must reply quickly and do its slow work afterwards. Keep polling as a fallback, because delivery is an optimization and not the only recovery path. See Webhooks.
Sources
Related posts
More in Integrations
- Cloud Tasks settings for Sume video jobs: pace submits, retry 429
Cloud Tasks defaults to 500 dispatches a second, far above a Pro plan's 24 accepted jobs. Pace submits and retry 429 queue_full with the same Idempotency-Key.
- Can DeepSeek V4.1 Flash call Sume's hosted MCP tools?
DeepSeek V4.1 Flash supports tool calls. Your code maps each call to Sume MCP tools_list, tools_schema and a gated paid call. The model never holds your key.
- Which account pays: fal's Active MCP account vs a Sume key
fal's MCP sends credits to the Active MCP account you choose. On Sume the API key or OAuth session selects the workspace, and no tool takes a workspace id.
- Ghost audio card with a Sume track: mp3, wav or ogg from desktop
Ghost's audio card takes .mp3, .wav and .ogg uploaded from the desktop editor, up to 1 GB by plan. Download the Sume artifact, then upload the file.
Written by Sume