Go errgroup SetLimit: submit Sume jobs within your concurrency budget
Cap in-flight Sume submissions in Go with errgroup.SetLimit, size the width from your concurrency limit, and keep one Idempotency-Key per item. 30 lines.

In Go, errgroup.Group.SetLimit(n) is the simplest way to keep at most n Sume submissions in flight: Go blocks when the limit is reached, so a plain loop becomes a bounded fan-out. Pick n from your effective concurrency_limit, not from a number you remember from the plan table.
The errgroup facts here are from the package page on pkg.go.dev (read 2026-10-10): SetLimit caps active goroutines, Go blocks while the limit is reached, and WithContext cancels the derived context the first time a function returns a non-nil error or Wait returns. The Sume facts are from Generation admission.
Where does the width come from?
Sume accepts valid jobs as queued while queue capacity remains, so the API will not stop you from firing 100 at once. The docs still recommend a client-side pace: new in-flight work is max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs), capped by queue_capacity_remaining. Submit responses carry those numbers in generation_limits when Sume can compute the snapshot.
| Plan | Processing concurrency | Queue capacity | Accepted capacity |
|---|---|---|---|
| Free | 1 | 5 | 6 |
| Pro | 4 | 20 | 24 |
| Startup | 8 | 40 | 48 |
| Scale | 20 | 100 | 120 |
The program
The sample reads the width from SUME_WIDTH, which you set from the dashboard Concurrency tab or from the last generation_limits you saw. Each item gets its own Idempotency-Key, so a retried run of the same batch returns the original jobs instead of billing new ones. On the first error, WithContext cancels the shared context and the remaining requests stop; jobs that were already accepted keep running and billing, so store their ids before you exit.
package main; import ("bytes"; "context"; "encoding/json"; "fmt"; "net/http"; "os"; "strconv"
"golang.org/x/sync/errgroup")
func submit(ctx context.Context, key, prompt string) (string, error) {
b, _ := json.Marshal(map[string]string{"prompt": prompt, "mode": "async"})
req, _ := http.NewRequestWithContext(ctx, "POST",
"https://api.sume.com/v1/image-1.0/generate", bytes.NewReader(b))
req.Header = http.Header{"X-Api-Key": {os.Getenv("SUME_API_KEY")},
"Content-Type": {"application/json"}, "Idempotency-Key": {key}}
res, err := http.DefaultClient.Do(req)
if err != nil { return "", err }
defer res.Body.Close()
var out struct{ Data struct{ RequestID string `json:"request_id"` } }
if err = json.NewDecoder(res.Body).Decode(&out); err != nil || res.StatusCode >= 300 {
return "", fmt.Errorf("submit %s: status %d: %v", key, res.StatusCode, err)
}
return out.Data.RequestID, nil
}
func main() {
width, _ := strconv.Atoi(os.Getenv("SUME_WIDTH"))
g, ctx := errgroup.WithContext(context.Background())
g.SetLimit(max(1, width))
ids := make([]string, 6)
for i := range ids {
g.Go(func() (err error) {
ids[i], err = submit(ctx, fmt.Sprintf("hero-batch-7-item-%d", i), "matte bottle on marble")
return
})
}
fmt.Println(g.Wait(), ids)
}What to watch for
Go 1.22 and later gives each loop iteration its own i, which is why the closure above is safe without a copy. The max builtin needs Go 1.21 or later. I compiled and vetted the program with Go before writing this page, but I did not run it against the live API.
If a submit returns 429 queue_full, the workspace has no accepted capacity left. Do not widen the group; wait for jobs to finish or cancel the queued ones, then retry with the same key. A 429 rate_limited is a separate event and retry-after carries the backoff.
- Width is a client pace, not a server limit; the API can still queue more.
- One idempotency key per item, stable across reruns.
- Persist ids as they arrive, not after
Waitreturns.
Sources
Related posts
More in Developers
- Grok Image on Sume: one image per call, so fan out four in Python
x-ai/grok-image lists n as 1 to 1 in the Sume catalog. How to get four variants with four parallel calls, what it costs, and how to stay under queue limits.
- Grok Imagine ignores aspect_ratio on image-to-video; Sume rejects it
xAI says image-to-video output matches the input image and ignores aspect_ratio. Sume's Grok row goes further and rejects the field. Crop the still first.
- Grok Imagine Lite's 10 requests per second vs Sume plan concurrency
xAI lists a 10 requests per second limit and Batch API for Grok Imagine 1.5 Lite. On Sume, a clip batch is bounded by plan concurrency and queue capacity.
- How many Format runs per minute can my Sume plan start?
Sume limits writes per minute by plan, from 120 on Free to 1200 on Scale, with reads at 40 times that. Read the rate-limit headers and back off on 429.
Written by Sume