Envoy's 15 s route timeout vs a 30 s Sume sync call
Envoy's RouteAction timeout defaults to 15s. Sume sync mode can wait 30 s. Behind Envoy or Istio, use async, or raise the route timeout above the sync wait.

Why does a gateway cut a sync call at 15 seconds?
Envoy's route configuration says RouteAction.timeout defaults to 15 seconds, 0 disables it, and the value includes retries. idle_timeout is a separate setting. Sume's sync mode waits at most 30 seconds before it returns a job envelope, so an Envoy route with defaults can give the client a 504 first.
| Setting | Value |
|---|---|
| Envoy RouteAction.timeout default | 15 seconds, includes retries |
| Envoy 0 | Disables the route timeout |
| Sume sync wait | At most 30 seconds |
| Sume async (default) | 202 with job envelope |
What should I change?
The cleanest fix is not to wait. Leave mode unset so Sume answers with a 202 envelope in well under a second, then follow status_url from the client or take a webhook. If you must wait synchronously, set the route timeout above 30 seconds and account for retries: a retry inside the route window can replay your POST, so keep the Idempotency-Key stable.
Setting the timeout to 0 removes the safety net for every request on that route, so prefer a specific longer value.
routes:
- match: { prefix: "/v1/avatar-1.0/" }
route:
cluster: sume_api
timeout: 45sWhat does the client do on a timeout?
Do not send the paid POST again. Poll the job you already created; the jobs docs say not to resubmit the original paid request only because a local process timed out. The Heroku variant of the same problem is in the H12 post.
Sources
Related posts
More in Developers
- Excel VBA: submit an AI video job and poll it with Application.OnTime
An Excel macro posts a Wan 3.0 job to Sume with ServerXMLHTTP, saves the job id in a cell and re-checks status with Application.OnTime without blocking Excel.
- Express receiver for Sume TTS job webhooks: raw body and HMAC
A Node Express route that verifies Sume job webhooks over the raw body, accepts rotated secrets, rejects an empty secret and answers 204 before any work.
- Fall back to a second image model after three 502s: a Python breaker
Sume returns 502 when an image job fails inside the wait budget. Count them, switch to a second model id after three, and use a new idempotency key per model.
- Gemini CLI and the hosted Sume server: do not rely on env in headers
Add the hosted Sume server to Gemini CLI with httpUrl and a bearer header. Gemini expands env vars only in the env block; set a timeout above jobs_wait.
Written by Sume