Go net/http: a Sume video submit retry loop with one Idempotency-Key
A 30-line Go function using net/http that retries a Sume video submit on 429 and 503, reads Retry-After, and reuses one Idempotency-Key. Run on a stub.

In Go, retry a Sume video submit by building a fresh http.Request on each attempt, setting the same Idempotency-Key header every time, and sleeping for Retry-After seconds on a 429 or 503 before the next try. The function below does that with the standard library only, in 30 lines. Build the request inside the loop, because a request body reader is consumed by the first send.
The same key is what makes this safe. Sume returns the original job for a replay, so if the first attempt reached the server and only the answer was lost, the retry does not create or bill a second job.
The code
Run it with SUME_API_KEY set. SUME_BASE is an override for tests and defaults to the production host. The client timeout is 35 seconds, which is above the 30-second cap on a synchronous wait. The loop returns the status code of the first answer that is not 429 or 503, so the caller decides what a 202 or a 400 means.
package main
import ("bytes"; "fmt"; "net/http"; "os"; "strconv"; "time")
func submit(base, key string, body []byte) (int, error) {
client := &http.Client{Timeout: 35 * time.Second}
for n := 0; n < 4; n++ {
req, _ := http.NewRequest("POST", base+"/v1/videos", bytes.NewReader(body))
req.Header.Set("Authorization", "Bearer "+os.Getenv("SUME_API_KEY"))
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", key)
res, err := client.Do(req)
if err != nil { return 0, err }
res.Body.Close()
if res.StatusCode != 429 && res.StatusCode != 503 { return res.StatusCode, nil }
wait := time.Duration(1<<n) * time.Second
if s, err := strconv.Atoi(res.Header.Get("Retry-After")); err == nil {
wait = time.Duration(s) * time.Second
}
fmt.Println(res.StatusCode, "waiting", wait)
time.Sleep(wait)
}
return 0, fmt.Errorf("still busy after 4 attempts")
}
func main() {
base := os.Getenv("SUME_BASE")
if base == "" { base = "https://api.sume.com" }
fmt.Println(submit(base, "mug-go-v2", []byte(`{"model":"seedance-2.5","prompt":"A mug"}`)))
}Details that matter in Go
- Close
res.Bodyon every path. An unclosed body keeps the connection from being reused. - Build the request in the loop.
bytes.NewReaderis read once, so reusing one request would send an empty body on the second try. strconv.Atoireads whole seconds only. ARetry-AfterHTTP date makes it fail, and the code falls back to 1, 2, 4 and 8 seconds.- A transport error returns at once. To retry it,
continueinstead, and keep the same key.
Status codes in this loop
The function does not read the response body, to stay short. In real code, decode the error object on a non-202 answer and log its request_id, because that is what support asks for first.
| Status | Loop action | Why |
|---|---|---|
| 202 | Return | The job was accepted |
| 429 | Wait and retry | rate_limited or queue_full; retry with the same key |
| 503 | Wait and retry | provider_capacity_exceeded; retry later with the same key |
| 400, 402, 404 | Return | A repeat would get the same answer |
| 409 | Return | idempotency_conflict: the key was used for another payload |
Proof it ran
The function ran against a local stub that sends 429 with Retry-After: 1, then 503, then 202 for a single key. The output was the two waiting lines (1s and 2s) and then 202 <nil>. Nothing was submitted to Sume.
If you need a pool of workers, give each one its own key per job and size the pool from your plan's accepted capacity, not from the number of CPU cores.
Using it in a worker
Wrap the function in a worker that takes jobs from a channel. Each job carries its own key, built from your own record, and the worker calls submit with that key. Because the key travels with the job, a restart of the worker replays the same keys and finds the accepted jobs instead of creating new ones.
Size the number of workers from your plan's accepted capacity. On a plan that accepts six jobs, ten workers will only produce queue_full answers. The submit response carries a generation_limits object with the remaining queue capacity, and it is a better input than a constant.
Add a context to the request with http.NewRequestWithContext if the worker must stop on shutdown. A canceled context ends the wait, and the job id of anything already accepted stays in your table.
Sources
Related posts
More in Developers
- Google made Omni Flash its default video model: pin it or sume/auto?
Google's docs now say to use Gemini Omni Flash as the default video model. On Sume you can pin gemini-omni-flash-1.1 or send sume/auto. What each one fixes.
- Google says Imagen is shut down: test your Imagen call on Sume now
Google's Imagen page says Imagen models are shut down. Sume lists Imagen 4 Fast and Ultra separately. One catalog check and one request tell you if yours works.
- GPT Image 2.5 inpainting with mask_url on Sume: steps and cost
Edit part of an image with mask_url and up to 16 references on openai/gpt-image-2.5. Billed about $0.066 an image on Sume. Steps and the limits.
- gpt-image-2.5 quality auto reserves max: holds from $0.22 to $0.89
On Sume, gpt-image-2.5 with quality auto and auto size reserves $0.8895 per image, while omitting quality reserves $0.2224. Hold table and the safe request.
Written by Sume