Submit Sume jobs from Go with a bounded pool and idempotency keys
Bound a Go worker pool to your workspace in-flight headroom, send a stable Idempotency-Key per item, and stop treating a full queue as a crash. Stdlib only.

Size the pool to your in-flight headroom, not to the number of items. Headroom is max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs), capped at queue_capacity_remaining. Give every item a stable Idempotency-Key that you derive from your own data, and a retry after a network error cannot create a second paid job. Sume publishes no Go SDK, so use net/http or a client generated from the OpenAPI document.
What the limits mean for a Go pool
Go makes it easy to launch thousands of goroutines, so the limit has to be explicit. The pool below is a channel with a fixed number of slots, which is the usual way to cap work without a library. Each goroutine takes a slot before it starts, and gives the slot back when it finishes.
Concurrency is a dispatch limit in Sume, not a submit limit. A valid job is accepted as queued while queue capacity remains, and only then moves to processing. A Go program that fans out 500 goroutines will not crash the API. It will fill your queue, and the 429 queue_full response is the signal that you went past it.
| Plan | Processing concurrency | Default queue capacity | Accepted capacity |
|---|---|---|---|
| Free | 1 | 5 | 6 |
| Pro | 4 | 20 | 24 |
| Startup | 8 | 40 | 48 |
| Scale | 20 | 100 | 120 |
| Enterprise | 20 | 100 | 120 |
Read the live numbers
Prefer the effective generation_limits.concurrency_limit on a submit response to this static table. Admin overrides can raise it, and the wave_size_hint field is only a hint for how big a submission wave can be. Never use it as the pool width.
The pool
The sample uses a buffered channel as a semaphore and a WaitGroup. The key is built from the batch name and the item id, so a re-run of the same batch replays the same keys.
package main
import ("fmt"; "net/http"; "os"; "strings"; "sync")
func submit(key, body string) (int, error) {
u := "https://api.sume.com/v1/image-1.0/generate"
req, _ := http.NewRequest("POST", u, strings.NewReader(body))
req.Header.Set("x-api-key", os.Getenv("SUME_API_KEY"))
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", key)
res, err := http.DefaultClient.Do(req)
if err != nil { return 0, err }
defer res.Body.Close()
return res.StatusCode, nil
}
func main() {
sem, wg := make(chan struct{}, 4), sync.WaitGroup{}
for _, id := range []string{"a", "b", "c"} {
wg.Add(1); sem <- struct{}{}
go func(id string) {
defer wg.Done(); defer func() { <-sem }()
body := `{"prompt":"hero shot ` + id + `","mode":"async"}`
code, err := submit("batch-001-"+id, body)
fmt.Println(id, code, err)
}(id)
}
wg.Wait()
}
Handle each answer
Handle the answers by code. Read the status of each response before you read its body. A 202 with a job_id means store it and poll later. A 429 rate_limited carries retry-after, so back off for that long. A 429 queue_full means wait for jobs to finish and then retry with the same key. A 402 insufficient_credits is not retryable, because the balance cannot cover the reservation. A 409 idempotency_conflict means you reused a key for a different body, so fix the key derivation, do not retry.
Keep the job id of every accepted item, and poll with backoff. A timeout in your program is not a failed job, and replaying the original request without the same key can charge twice.
Habits for a long batch
Three habits keep a batch honest. Each one is cheap to add before the first run, and expensive to fix after a duplicate charge. First, derive the key from stable ids such as the batch name and the item, never from a random value made per attempt, because a random key defeats the whole purpose. Second, change the key when you change the body on purpose, because the same key with a different body is a 409. Third, refresh your view of the limits before each wave. The counts in generation_limits are a snapshot, and workers or other clients can change them right after the response.
Choosing the width
For the width itself, start with the effective concurrency_limit from your last submit response. For example, a Pro workspace with no active or queued jobs has headroom of 4. With 3 processing and 1 queued, the headroom is 0, so wait and look again. Because the count is a snapshot, count every job that you just submitted against the budget until you fetch a fresh one. The channel in the sample is fixed at 4 to keep the code short, and a real program should reset it from the live numbers.
Sources
Related posts
More in Developers
- Sume 400 unsupported_capability on 4:5: fall back to 3:4 in Python
A 4:5 request on a Sume video model returns 400 unsupported_capability before billing. Catch it in Python, resubmit at 3:4, then crop to Meta's 4:5 Feed ratio.
- Sume 429: read error.details.scope, and rate_limit_unavailable
A Sume 429 names the budget in error.details.scope and gives retry_after_seconds. A degraded rate_limit_unavailable 429 is a different case. Python handler.
- Sume schedule run 403: a service-account key cannot start action runs
Starting a Sume schedule run with a service-account key fails with 403 insufficient_scope and service_account_action_runs_unsupported. Use a user API key.
- Why sume/auto returns 400 for 21:9, not a Seedance route
sume/auto accepts 3 to 10 s in 16:9 or 9:16. Ask for 21:9 or 15 s and you get 400 unsupported_capability, not a quiet reroute to Seedance.
Written by Sume