Batch of 12 H3 Max Recast swaps: queue limits by Sume plan

Twelve Recast swaps fit in one submit on Pro and above (24+ accepted jobs); on Free only 6 are accepted at once. The queue table and the $45 reserve at 768p.

5 min readSume
All posts

Twelve H3 Max Recast swaps can be submitted in one go on the Pro plan or above, because Pro accepts 24 paid generation jobs at a time (4 processing, 20 queued). On Free, Sume accepts 6 at once (1 processing, 5 queued), so a batch of 12 needs two waves or it returns 429 queue_full. Ten second swaps at 768p reserve $3.75 each, $45.00 for the batch.

This is the shape of a typical Q4 job: one winning clip, a dozen presenters, one deadline.

Concurrency is a dispatch limit, not a submit limit

Sume's admission docs separate four controls: generation concurrency (jobs in processing), queue capacity (jobs accepted but waiting), submit rate limits, and the spendable balance. A full concurrency limit does not reject work. Valid jobs wait in queued until a slot opens, and the submit only fails with queue_full when the queue is full too (Generation admission).

Sume generation admission by plan, default limits (read 2026-10-03)
PlanProcessingQueueAccepted at onceFits 12 swaps
Free156no, 6 at a time
Pro42024yes
Startup84048yes
Scale20100120yes

What a batch of 12 looks like

On Pro, all twelve submits return queued or processing. Four run; eight wait; each starts as a slot frees. Recast jobs are paid generation jobs like any other, so the same rules apply: store each job_id, poll with backoff, and fetch results at status: completed.

On Free, submit six, then wait for at least one to reach a terminal state before the next. Do not treat queue_full as a failed job. Retry the same request with the same Idempotency-Key once capacity opens.

The docs tell you to prefer the effective concurrency_limit in the submit response over the static table, since admin overrides can raise it. Read generation_limits from the first response and size the next wave from queue_capacity_remaining.

  • Each accepted job reserves its estimated price at submit. Twelve 10 second swaps at 768p reserve $45.00 in total; at 1080p, $67.50.
  • If the balance cannot cover a reservation, that submit fails with 402 insufficient_credits before any provider work starts.
  • Queued jobs can be canceled before they start; once generation has begun, cancel returns 409 job_generation_already_started.

A submit loop that respects both limits

Keep the loop boring: submit while the response says there is room, sleep when it does not, and always send an idempotency key derived from the batch and the item so a retry cannot double-charge. Use a webhook for completion if you do not want a poller; the signature must be verified before you download anything (Webhooks).

Pacing a batch in practice

A batch is a loop that submits and then waits, not a loop that submits and sleeps. Submit jobs with distinct idempotency keys, record each job id, and let queued jobs wait; a full processing limit is not an error. Only a full queue returns 429 queue_full, and that is your signal to pause submitting.

Prefer a webhook to polling for a large batch, since terminal events tell you when a slot has freed. Keep a status poll as a fallback for missed deliveries. If the batch mixes resolutions, the job count is what the queue sees, not the dollar amount, so twelve cheap jobs and twelve expensive ones occupy the same capacity.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume