Free plan 120 writes a minute: a 100-item bulk queue is one request
A Sume bulk create is one POST against the write budget, and polling spends the separate read budget. Here are the per-plan limits and what that means.

A Sume bulk create is one POST and spends one request from the write budget, whatever the item count. On the Free plan that budget is 120 writes a minute, so one request carrying 100 items uses 1 of them. Polling the queue draws on a separate read budget that is forty times larger.
The per-plan numbers
Each key has a per-minute request budget across /v1. A read is any GET; a write is every other request, including run and queue creates, cancel and webhook redeliver. A 429 names the budget in error.details.scope and adds retry-after.
| Plan | Writes | Reads |
|---|---|---|
| Free | 120 | 4800 |
| Pro | 300 | 12000 |
| Startup | 600 | 24000 |
| Scale | 1200 | 48000 |
| Enterprise | Contracted (Scale until provisioned) | Contracted |
What this changes in practice
The docs count requests, so sending 100 single POST .../runs calls spends 100 writes, while one POST .../bulk-runs spends one. The rate budget is rarely your constraint in a bulk run. Workspace generation concurrency, the queue window of 1 to 16, and spend caps are the limits that shape the batch.
- Poll the queue with backoff; a poll that returns
429or503is transient and the queue keeps draining. - Each response carries
ratelimit-limit,ratelimit-remainingandratelimit-reset, so log them rather than guess. - Cancelling children spends writes too, one per cancelled run.
Sizing a batch against the budget
Imagine a Free workspace with a 100-item batch. Sent as single creates, that is 100 of the 120 writes available in a minute, leaving 20 for cancels, redelivers and retries. Sent as one queue, it is one write, and the remaining 119 stay free for everything else.
Polling is the other half of the arithmetic. The read budget is forty times the write budget, so a poll every ten seconds on a single queue uses a tiny share of even the Free plan's reads. The point is to poll with backoff because video children take minutes, not because the limit is tight.
If you hit a 429 anyway, read error.details.scope before you change anything. A read scope means your polling is too eager; a write scope means creates, cancels or redelivers are bunching together, and the fix is to spread them or move to a plan with a larger budget.
Tradeoff
The docs state the budget in requests, not in runs. They do not say whether child runs started by a queue count against your key, so do not plan on them being free or on them being charged; watch the headers during a pilot of three rows.
Sources
Related posts
More in Formats
- Localize one Sume Format into six languages: one bulk item each
To make one ad in six languages, send six items to one Sume bulk queue with the language in input and one Idempotency-Key scheme. Here is the shape and limits.
- Log format.version on every bulk child: which recipe made which video
A Sume bulk queue can span a Format edit. Store format.version from each child receipt so you know which recipe made which video.
- Make a music video from your own footage with AI: Sume restyle
sume-restyle redoes a clip you own with the same motion and cuts, in chunks of up to 30 seconds. What to supply, what stays, and the audio rule.
- No queue webhook on Sume bulk runs: count item webhooks instead
Sume bulk queues have no queue-level webhook. Put a webhook_url on each item and count terminal events to know a season is finished.
Written by Sume