nginx proxy_read_timeout is 60s, but Sume gives up at 10s

nginx waits up to 60 seconds for your app by default, yet a Sume webhook attempt times out at 10. Why raising proxy timeouts does not help and what to do.

4 min readSume
All posts

Raising nginx's proxy timeouts will not fix a slow Sume webhook receiver, because the clock that matters is Sume's. nginx's proxy module documents proxy_read_timeout with a default of 60 seconds (read 2026-10-10), but each Sume delivery attempt is cut off after 10 seconds, so your proxy can be patiently waiting on an upstream that Sume has already abandoned.

The fix is on the application side: acknowledge the delivery fast and do the slow work afterward. The rest of this post shows where each timeout sits, so you can tell which one produced the failure you are looking at.

Two clocks, different owners

A delivery crosses three hops: Sume's delivery worker, your nginx, and your application. Sume's webhooks page sets the sender-side rules: a delivery gets up to 10 attempts, 30 seconds apart, and each attempt times out after 10 seconds. A 2xx response counts as delivered.

nginx has its own timers for the hop behind it. The proxy module page lists proxy_connect_timeout, proxy_send_timeout and proxy_read_timeout, each defaulting to 60 seconds, and says the read timeout applies between two successive read operations, not to the whole response (read 2026-10-10). That makes 60 seconds a ceiling for a silent upstream, not a promise that a 40-second handler will finish in time for the caller.

Timeouts on a Sume webhook path (nginx defaults from its proxy module page, read 2026-10-10)
HopSettingValueWho controls it
Sume to your edgePer-attempt timeout10 sSume, fixed
Sume to your edgeAttempts and spacing10 attempts, 30 s apartSume, fixed
nginx to your appproxy_connect_timeout60 s defaultYou
nginx to your appproxy_send_timeout60 s defaultYou
nginx to your appproxy_read_timeout60 s default, between two readsYou

What a slow handler looks like from each side

Suppose your handler downloads the generated file before it responds, and that takes 25 seconds. nginx keeps the connection open because the upstream is still within its 60-second allowance. Sume closes the attempt at 10 seconds and records a timeout, then tries again 30 seconds later. Your access log can show a request that eventually returns 200 or 499, while Sume's delivery status goes to retrying and, after ten misses, to exhausted.

The mismatch is the signature of this bug: the logs on your side look healthy, and the sender's look broken. Sume's errors and credits page lists the delivery states pending, delivering, delivered, retrying, failed and exhausted, and the receipt records last_status_code and last_error, so a timeout shows up there as a transport error rather than as your body.

The fix: acknowledge, then work

Shorten the work in the request, not the proxy's patience. A receiver that verifies the signature, stores the raw event, returns 200 and hands the rest to a queue stays far under 10 seconds regardless of file size. Because retries are possible, key the stored event on job_id, which Sume's docs say to treat as the idempotency key.

Leave the nginx defaults alone unless you have another reason to change them. Lowering proxy_read_timeout on this one location to a few seconds is a reasonable guard, since it makes nginx return an error to Sume instead of hanging past a point where Sume has already left.

  • Return 2xx only after the event is durably stored, not after downstream work.
  • Verify the signature before enqueueing, and refuse to run with an empty secret.
  • Do the file copy, transcoding or database fan-out from the queue worker.
  • Use a separate location block for the webhook path so its timeouts do not affect other routes.

Recovering deliveries that already ran out

If a slow release already burned the ten attempts, the job itself is untouched. Fix the receiver, then call POST /v1/jobs/{job_id}/webhook/redeliver with a key holding jobs:write. Sume re-sends the job's real terminal event with a fresh timestamp and signature, and the docs say this works after the automatic attempts are used up and does not count as one of them.

Redeliver does not change the destination. If the URL was wrong, that is a new job. For a handful of misses, redeliver one by one; for a larger gap, poll the job status URL for the affected jobs, which Sume's docs describe as the recovery path for events that never arrived.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume