Load-test Sume job polling with k6: reads have their own budget

A k6 script that polls one finished Sume job from 5 virtual users, and what 4,800 reads a minute on Free means for any 429 you see.

4 min readSume
All posts

What the test should prove

Before you ship a poller that watches dozens of 30-second video jobs, you want to know that your own polling will not trip a 429. Sume answers that with a split budget: each API key gets a per-minute request budget for all of /v1, and reads and writes are counted separately. A read is any GET or HEAD, such as a poll of status_url, events_url or result_url.

That split is why a status loop cannot starve your submits. It is also why you should load-test reads against a job that is already finished: the status endpoint answers at once, nothing waits, and you measure only your own request rate.

Reads and writes per minute by plan (read 2026-10-05)

Per-key request budgets from the Sume authentication page, read 2026-10-05
PlanWrites per minuteReads per minute
Free1204800
Pro30012000
Startup60024000
Scale120048000

The k6 script

Set SUME_API_KEY and JOB_ID of a completed job, then run k6 run poll.js. Five virtual users each send one read every two seconds, which is about 150 reads a minute, far below the Free budget. The check accepts 200 and counts 429 separately so you can see when the ceiling arrives as you raise the user count.

import http from "k6/http";
import { check, sleep } from "k6";

export const options = { vus: 5, duration: "1m" };
const params = { headers: { Authorization: "Bearer " + __ENV.SUME_API_KEY } };
const url = "https://api.sume.com/v1/jobs/" + __ENV.JOB_ID + "/status";

export default function () {
  const res = http.get(url, params);
  check(res, {
    "status 200": (r) => r.status === 200,
    "not rate limited": (r) => r.status !== 429,
  });
  sleep(2);
}

Reading the outcome

  • A 429 body names the budget in error.details.scope, read or write, so a read 429 never means your submits were throttled.
  • When a 429 arrives, wait for the retry-after header; do not retry in a tight loop.
  • Do not hard-code 4800. The ratelimit-limit header on each response is the authority for the deployment you call.
  • Raise vus in steps (5, 20, 50) and stop at the first 429; that is your measured ceiling for one key.
  • Request rate is not generation capacity: polling faster never lets more jobs run at once.

Keep the real poller on the server-suggested interval. When the status envelope carries next_poll_after_seconds, sleep that long, and use exponential backoff when it is absent.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume