Sume write limits as submits per second: a Python pacer by plan

Free 120, Pro 300, Startup 600, Scale 1200 writes per minute is 2, 5, 10 and 20 per second. A fixed-window pacer in Python that never trips the 429.

4 min readSume
All posts

The Sume rate limit is documented per minute, but a batch job thinks in submits per second. The write budget of each plan converts directly: Free 120 per minute is 2 per second, Pro 300 is 5, Startup 600 is 10 and Scale 1200 is 20. Polling does not count against it, because reads have a separate bucket at forty times the write number.

The conversion

Write budget by plan, from the Sume docs (read 2026-10-04)
PlanWrites per minutePer secondReads per minute
Free12024800
Pro300512000
Startup6001024000
Scale12002048000

What counts as a write

Any request that is not a GET or HEAD is a write: run creation, cancellation and uploads. Two POSTs that submit nothing are reads: /v1/generation/admission-preview and the MCP endpoint itself. A webhook redelivery is a POST as well, so it spends from the same write budget.

A pacer

Because the window is a fixed 60 seconds, you can burst up to the plan number at the start of a window and then wait. A steady rate is usually friendlier to the generation queue, which has its own limit of concurrent jobs. This pacer spreads submits evenly over the window and does not need to know the server clock.

import time

class Pacer:
    def __init__(self, per_minute, clock=time.monotonic, sleep=time.sleep):
        self.gap = 60.0 / per_minute
        self.clock, self.sleep = clock, sleep
        self.next_at = clock()

    def wait(self):
        now = self.clock()
        if now < self.next_at:
            self.sleep(self.next_at - now)
        self.next_at = max(now, self.next_at) + self.gap

pacer = Pacer(per_minute=300, sleep=lambda s: None)  # Pro: one submit every 0.2 s
for _ in range(3):
    pacer.wait()
print(round(pacer.gap, 3))  # 0.2

Leave some headroom

Set the pacer a little under the plan number, for example 90 percent, if anything else uses the same key or the same workspace owner. All of those requests share the account bucket. Request rate also does not change the concurrency limit: the generation_limits object reports how many jobs can run at once.

A worked example

Take a batch of 600 submits. On Free, at 120 a minute, it needs five minutes. On Pro, at 300, it needs two minutes. On Startup, at 600, one minute. On Scale, at 1200, 30 seconds. Those are floors from the request limit alone. Real runtime is longer because each job also waits for generation capacity, which the plan's concurrency limit controls separately.

So plan for the slower of the two limits. Compute the submit time from the table, then read generation_limits and size waves so that no more jobs are in flight than the plan allows. A queue_full 429 means the second limit was hit, and the pacer cannot prevent it.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume