Load-test Sume job polling with k6: reads have their own budget
A k6 script that polls one finished Sume job from 5 virtual users, and what 4,800 reads a minute on Free means for any 429 you see.

What the test should prove
Before you ship a poller that watches dozens of 30-second video jobs, you want to know that your own polling will not trip a 429. Sume answers that with a split budget: each API key gets a per-minute request budget for all of /v1, and reads and writes are counted separately. A read is any GET or HEAD, such as a poll of status_url, events_url or result_url.
That split is why a status loop cannot starve your submits. It is also why you should load-test reads against a job that is already finished: the status endpoint answers at once, nothing waits, and you measure only your own request rate.
Reads and writes per minute by plan (read 2026-10-05)
| Plan | Writes per minute | Reads per minute |
|---|---|---|
| Free | 120 | 4800 |
| Pro | 300 | 12000 |
| Startup | 600 | 24000 |
| Scale | 1200 | 48000 |
The k6 script
Set SUME_API_KEY and JOB_ID of a completed job, then run k6 run poll.js. Five virtual users each send one read every two seconds, which is about 150 reads a minute, far below the Free budget. The check accepts 200 and counts 429 separately so you can see when the ceiling arrives as you raise the user count.
import http from "k6/http";
import { check, sleep } from "k6";
export const options = { vus: 5, duration: "1m" };
const params = { headers: { Authorization: "Bearer " + __ENV.SUME_API_KEY } };
const url = "https://api.sume.com/v1/jobs/" + __ENV.JOB_ID + "/status";
export default function () {
const res = http.get(url, params);
check(res, {
"status 200": (r) => r.status === 200,
"not rate limited": (r) => r.status !== 429,
});
sleep(2);
}Reading the outcome
- A 429 body names the budget in error.details.scope, read or write, so a read 429 never means your submits were throttled.
- When a 429 arrives, wait for the retry-after header; do not retry in a tight loop.
- Do not hard-code 4800. The ratelimit-limit header on each response is the authority for the deployment you call.
- Raise vus in steps (5, 20, 50) and stop at the first 429; that is your measured ceiling for one key.
- Request rate is not generation capacity: polling faster never lets more jobs run at once.
Keep the real poller on the server-suggested interval. When the status envelope carries next_poll_after_seconds, sleep that long, and use exponential backoff when it is absent.
Sources
Related posts
More in Developers
- Log usage.cost for every Sume video job in Python and total a batch
A finished Sume video job returns usage.cost in USD. Log it with the job id and model, and sum it for a batch. Here is a short Python script that does it.
- Log usage.cost from the Sume images response to CSV in Python
POST /v1/images returns usage.cost as the billed USD amount. A Python logger that writes model, image count and cost per image to a CSV for budget reviews.
- Longest AI video Sume can render: 1800 s from 60 clips
One Timeline render caps at 1800 seconds. With 30-second clips that is 60 clips, with 200 slots allowed. The math and the render price follow.
- Lyria 3.5 prompts for video BGM: section tags and timestamps
Use [Verse]/[Chorus]/[Bridge] tags and [0:00-0:10] timestamps in a Lyria 3.5 prompt to pin an arc to a video. A worked prompt and a Sume call are below.
Written by Sume