Make removed Cycles per run on Oct 1: pace Sume submits instead
Make removed its Cycles per run setting on October 1, 2026. How to keep a Sume batch inside queue and rate limits without it, and when to use a bulk queue.

Make removed the Cycles per run setting on October 1, 2026, so a scenario that leaned on several cycles to space out paid calls needs its pacing somewhere else. With Sume the cleaner place is the Sume side: submit async with an Idempotency-Key, read generation_limits from each response, and stop adding work when headroom runs out.
Make's 2026 release notes (read 2026-10-02) say, in an entry dated August 13, 2026: "Make will remove the Cycles per run setting on October 1, 2026," and that users who employed more than one cycle in their scenarios must take action. The page content I read does not say what replaces the setting or how Make now behaves for a scenario that still had several cycles, so check your own scenarios in Make rather than assuming. This post only covers the Sume half: what keeps a batch safe when your scenario no longer paces itself.
Why does a batch of paid Sume calls need pacing at all?
Sume accepts paid generation jobs queue-first. A valid submit returns a durable job id even when the workspace is already running its maximum, and the job waits in queued (Generation admission). Concurrency is a dispatch limit, not a submit limit. Three other controls can still reject you: queue capacity (429 queue_full), submit rate limits (429 rate_limited), and balance (402 insufficient_credits).
So an unpaced scenario that fires every row at once will not melt Sume, but it can fill the queue, trip a rate limit, and leave half the batch failed with a plan-sized queue ahead of it. Pacing keeps the batch inside what your plan can actually process.
| Control | Error when full | What to do |
|---|---|---|
| Queue capacity | 429 queue_full | Wait for jobs to finish, retry with the same idempotency key |
| Submit rate limit | 429 rate_limited | Back off using retry-after when present |
| Balance | 402 insufficient_credits | Stop the batch; do not retry blindly |
| Concurrency | None (jobs wait as queued) | Treat queued as normal, not failure |
How do I size a wave from the submit response?
Each generation submit response carries generation_limits when Sume can compute it. The docs give a budget for new in-flight work: max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs), capped by queue_capacity_remaining. At zero headroom, wait and refresh before submitting more.
The same page warns about wave_size_hint: it is a submission-wave hint only, not a concurrency limit, and must not be used to size in-flight work. In Make that means an HTTP module that writes generation_limits into a data store or variable, plus a router that stops the loop when the budget is zero.
- Store the job id from every submit so a restart can resume from
status_url. - Count each newly submitted job against the budget until the next snapshot.
- Never lift the headroom number from
plan_concurrency_limit; use the effectiveconcurrency_limit.
When is a Format bulk queue the better replacement?
If each item is a saved Format run rather than a single model call, Sume can hold the pacing for you. POST /v1/formats/{handle}/{slug}/bulk-runs queues up to 100 runs with a concurrency window of 1 to 16, and you poll GET /v1/format-run-queues/{queue_id} for counts (Bulk runs). One HTTP module submits the whole list; a second polls the queue. There is no public cancel-queue endpoint, only a cancel per child run.
This needs an API key with formats:write and formats:read, and service-account keys cannot create Format runs or bulk queues. It does not help if you call a model route such as /v1/image-1.0/generate directly, because those are jobs, not runs.
curl -sS -X POST "https://api.sume.com/v1/formats/chase/product-promo/bulk-runs" \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: $(uuidgen)" \
-d '{"concurrency":3,"items":[{"instruction":"clip 1"},{"instruction":"clip 2"}]}'What should the retry path look like in a Make scenario?
Reuse one Idempotency-Key per item (for example batch-17-row-0042), so a retried HTTP module returns the original job instead of billing twice. Sume returns 409 idempotency_conflict if the same key is reused for a different payload, so derive the key from the row, not from a counter that shifts. For the error-route side, see Make.com error handling: retry a paid API call.
Do not resubmit just because a Make module timed out. The submit already returned a job id; poll GET /v1/jobs/{id}/status and honour next_poll_after_seconds. A local timeout does not cancel the job, which keeps running and billing.
What does Sume not do here?
Sume does not know about your Make cycles, and it does not slow submits down for you: the only brakes are the queue, the rate limit and the balance. It also offers no progress webhook; job webhooks fire on job.completed, job.failed and job.canceled only, so keep status polling as a backup.
Sources
Related posts
More in Integrations
- Cline MCP remote not connecting: set type streamableHttp for Sume
Cline treats a remote server with no type as legacy SSE. Sume's hosted MCP is streamable HTTP and answers GET /mcp with 405, so set type to streamableHttp.
- Mistral connectors: confirm Sume's paid tools before they run
Add Sume as a Mistral custom MCP connector, then use tool_configuration include and requires_confirmation to keep paid generation behind a human check.
- n8n 3.0 removed the HTTP Request Tool: call Sume from an agent
The legacy HTTP Request Tool is gone in n8n 3.0. Wire the HTTP Request node into the AI Agent Tool input to submit Sume jobs, or use hosted MCP.
- n8n 3.0 SSRF blocklist adds 100.64.0.0/10: Sume webhook paths
n8n 3.0 adds 100.64.0.0/10 to its SSRF blocklist. Sume is public HTTPS, so calls are fine, but webhook URLs must be public for Sume to reach them.
Written by Sume