Higgsfield concurrency limit returns 400, not 429: Sume's answer
Higgsfield answers 400 at your concurrency cap and sends no Retry-After. Sume queues valid jobs and returns 429 queue_full only when the queue is full.

When you hit your cap, Higgsfield returns 400 Bad Request with a message such as Maximum number of concurrent requests (4) has been reached, and its docs say the API publishes no standard rate-limit headers or Retry-After. Sume does not reject a valid job because concurrency is full. It accepts the job as queued, and returns 429 queue_full only when the queue itself is full.
What Higgsfield limits
Higgsfield defines its limit as concurrency: how many requests may be queued or processing at once. The number depends on the account and model, and you read it in the Higgsfield Console. A submit over the cap fails on the spot, so your client needs its own queue or semaphore, which is what the docs recommend.
What Sume limits instead
Sume separates four controls: processing concurrency, queue capacity, submit rate limits, and read-endpoint limits. Concurrency is a dispatch limit, not a submit limit. Queue capacity defaults to max(3, concurrency_limit x 5), and generation concurrency is plan-only: a prepaid top-up does not raise it.
| Situation | Higgsfield | Sume |
|---|---|---|
| Over the concurrency cap | 400 with a message; submit fails | Accepted as queued while queue capacity remains |
| Nothing left to queue into | Not applicable | 429 queue_full |
| Too many requests | Not described as a 429 case | 429 rate_limited |
| Retry hints | No Retry-After or rate-limit headers | ratelimit-* headers and retry-after can be present |
| Where to see the cap | Higgsfield Console | Dashboard Concurrency tab, generation_limits |
What your retry code should do
The practical difference is in your retry code. Against Higgsfield, a 400 on submit means wait for one of your own requests to finish. Against Sume, a 429 queue_full means the same, but a plain 429 rate_limited means back off using retry-after. In both cases, send an idempotency key so a retry cannot make a second paid job.
- Higgsfield: cap in-flight work with a semaphore sized to your console limit.
- Sume: read
generation_limitsand size submission waves from the remaining queue capacity. - Back off with jitter; do not poll in a tight loop on either API.
- Reuse the same
Idempotency-Keyon a retried submit.
Polling is separate
Both providers ask you to separate polling traffic from submission traffic. On Sume, read and status endpoints can be rate limited too, and the docs call that polling backpressure rather than generation concurrency.
Sources
Related posts
More in Developers
- Genjutsu missing from /v1/video-router/models: when it is listed
If higgsfield-genjutsu is not in GET /v1/video-router/models, Sume hides it when its provider is not configured. How to check, and what to use instead.
- Higgsfield hf_webhook query parameter vs Sume callback_url
Higgsfield takes the webhook as an hf_webhook query parameter; Sume takes callback_url in the body and signs the delivery. Payload shapes and a Python verifier.
- Higgsfield Idempotency-Key: 422 on a changed body, vs Sume 409
Both APIs replay the original job for a repeated Idempotency-Key. A changed body gets 422 on Higgsfield and 409 idempotency_conflict on Sume. Rules compared.
- Higgsfield output URLs last at least 7 days: copy files out
Higgsfield keeps generated output for at least seven days and may remove it later. Where Sume serves a finished video, and why to copy it to your own storage.
Written by Sume