Submit Sume jobs from Go with a bounded pool and idempotency keys

Bound a Go worker pool to your workspace in-flight headroom, send a stable Idempotency-Key per item, and stop treating a full queue as a crash. Stdlib only.

5 min readSume
All posts

Size the pool to your in-flight headroom, not to the number of items. Headroom is max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs), capped at queue_capacity_remaining. Give every item a stable Idempotency-Key that you derive from your own data, and a retry after a network error cannot create a second paid job. Sume publishes no Go SDK, so use net/http or a client generated from the OpenAPI document.

What the limits mean for a Go pool

Go makes it easy to launch thousands of goroutines, so the limit has to be explicit. The pool below is a channel with a fixed number of slots, which is the usual way to cap work without a library. Each goroutine takes a slot before it starts, and gives the slot back when it finishes.

Concurrency is a dispatch limit in Sume, not a submit limit. A valid job is accepted as queued while queue capacity remains, and only then moves to processing. A Go program that fans out 500 goroutines will not crash the API. It will fill your queue, and the 429 queue_full response is the signal that you went past it.

Generation limits by plan (read 2026-10-05)
PlanProcessing concurrencyDefault queue capacityAccepted capacity
Free156
Pro42024
Startup84048
Scale20100120
Enterprise20100120

Read the live numbers

Prefer the effective generation_limits.concurrency_limit on a submit response to this static table. Admin overrides can raise it, and the wave_size_hint field is only a hint for how big a submission wave can be. Never use it as the pool width.

The pool

The sample uses a buffered channel as a semaphore and a WaitGroup. The key is built from the batch name and the item id, so a re-run of the same batch replays the same keys.

package main

import ("fmt"; "net/http"; "os"; "strings"; "sync")

func submit(key, body string) (int, error) {
	u := "https://api.sume.com/v1/image-1.0/generate"
	req, _ := http.NewRequest("POST", u, strings.NewReader(body))
	req.Header.Set("x-api-key", os.Getenv("SUME_API_KEY"))
	req.Header.Set("Content-Type", "application/json")
	req.Header.Set("Idempotency-Key", key)
	res, err := http.DefaultClient.Do(req)
	if err != nil { return 0, err }
	defer res.Body.Close()
	return res.StatusCode, nil
}

func main() {
	sem, wg := make(chan struct{}, 4), sync.WaitGroup{}
	for _, id := range []string{"a", "b", "c"} {
		wg.Add(1); sem <- struct{}{}
		go func(id string) {
			defer wg.Done(); defer func() { <-sem }()
			body := `{"prompt":"hero shot ` + id + `","mode":"async"}`
			code, err := submit("batch-001-"+id, body)
			fmt.Println(id, code, err)
		}(id)
	}
	wg.Wait()
}

Handle each answer

Handle the answers by code. Read the status of each response before you read its body. A 202 with a job_id means store it and poll later. A 429 rate_limited carries retry-after, so back off for that long. A 429 queue_full means wait for jobs to finish and then retry with the same key. A 402 insufficient_credits is not retryable, because the balance cannot cover the reservation. A 409 idempotency_conflict means you reused a key for a different body, so fix the key derivation, do not retry.

Keep the job id of every accepted item, and poll with backoff. A timeout in your program is not a failed job, and replaying the original request without the same key can charge twice.

Habits for a long batch

Three habits keep a batch honest. Each one is cheap to add before the first run, and expensive to fix after a duplicate charge. First, derive the key from stable ids such as the batch name and the item, never from a random value made per attempt, because a random key defeats the whole purpose. Second, change the key when you change the body on purpose, because the same key with a different body is a 409. Third, refresh your view of the limits before each wave. The counts in generation_limits are a snapshot, and workers or other clients can change them right after the response.

Choosing the width

For the width itself, start with the effective concurrency_limit from your last submit response. For example, a Pro workspace with no active or queued jobs has headroom of 4. With 3 processing and 1 queued, the headroom is 0, so wait and look again. Because the count is a snapshot, count every job that you just submitted against the budget until you fetch a fresh one. The channel in the sample is fixed at 4 to keep the code short, and a real program should reset it from the live numbers.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume