Hackathon app on the Sume Free plan: 1 seat, 6 accepted jobs
A weekend demo on Free can have one job processing and five queued. How to design the UI, the retries and the demo script around that, with the real limits.

On the Free plan, one paid generation job can be processing at a time and five more can wait in queued, so at most six accepted jobs. Writes are limited to 120 a minute and reads to 4,800 a minute. A hackathon demo that fires ten video requests at once will see four of them answered 429 queue_full; one that fires three will see all three accepted and run them one after another.
That is workable if you design for it: queue on your side, show honest states, and rehearse with the numbers.
The limits that shape a demo
Values are from the generation admission and authentication pages. Queue capacity defaults to max(3, concurrency_limit x 5), which is 5 for a concurrency of 1.
| Limit | Value | Consequence for a demo |
|---|---|---|
| Processing concurrency | 1 | Jobs run one at a time |
| Queue capacity | 5 | Five more jobs can wait |
| Accepted jobs | 6 | The seventh submit gets 429 queue_full |
| Writes per minute | 120 | Not the binding limit |
| Reads per minute | 4800 | Poll freely, with backoff |
Design rules
Because Sume reports queue counts and not a queue position or ETA, do not promise one on screen. Show queued, processing and a finished clip, and let the user see the result as soon as it is ready.
Do not treat queued as a failure, and do not resubmit because your local worker timed out: an idempotency key makes a retry return the original job.
- Keep your own pending list and submit the next item only when
queue_capacity_remainingfrom the last response is above zero. - Pre-generate the clips you need for the demo script and cache the Sume artifact URLs; a judge should never wait for a render.
- Cancel is possible only before generation starts, so give users a cancel button only on
queueditems. - Prefer images for live interactions: they often finish inside the 30-second wait on
POST /v1/images, while video jobs usually do not.
When to leave Free
The plan, not a top-up, sets concurrency; the docs say prepaid top-ups do not raise it. If the demo turns into a pilot, Pro is 4 processing and 24 accepted. Read the effective generation_limits.concurrency_limit from a submit response instead of copying a table, because an admin override can change it.
A rehearsal checklist
Run the demo path twice the night before with a fresh workspace state. Check what happens at six jobs, then at seven, and confirm that your UI shows a calm message on 429 queue_full instead of a stack trace.
Also check credits: the first paid generation needs a balance, and 402 insufficient_credits is a different problem from a full queue, with a different fix.
Finally, keep a plain fallback. If a render fails during the live demo, show the cached clip and say so; judges forgive a disclosed fallback far more readily than a spinner that never ends. A failed job refunds its reserved credits, so a retry on a fresh key costs only the successful run.
Sources
Related posts
More in Developers
- GPT Image 1 to GPT Image 2.5 on Sume: what changes in the output
Moving from GPT Image 1 to ChatGPT Image 2.5 on Sume changes the response (URL, not base64), default quality, size grid and failures.
- What is a partial transcript in streaming speech to text?
A partial is a provisional transcript a streaming model revises as audio arrives. Why subtitles for a finished clip only need final text and word times.
- What to show a viewer while an avatar video job is queued
Avatar jobs on Sume are async: queued, processing, then completed, failed or canceled. A status-to-UI map for waiting screens, with polling rules.
- When is async TTS the right choice? Sync wait, poll or webhook
Async TTS is right for voiceovers, batches and anything a person is not watching a spinner for. Sume's sync wait stops at 30 seconds; Flash claims 45 ms.
Written by Sume