Kling API rate limit: concurrency by package and error 1303
Kling's API limits concurrent tasks by resource package, not requests per second. Over the cap, a create fails with HTTP 429, code 1303.

The Kling API's main limit is concurrency, not requests per second: it caps how many generation tasks your account can run at once, set by the resource packages you've bought, and Kling's concurrency page says it imposes no QPS limit. When you're at that cap, a new create call fails with HTTP 429 and error code 1303 until a running task ends.
These facts come from Kling's own Concurrency Rules, video pricing, Unit Deduction Detail, Quick Start and error codes pages, read on 2026-09-29. Those pages don't publish a concurrency number per package, so read yours from the package you buy.
How does Kling API concurrency work?
Only the task creation call counts. Query calls don't consume concurrency. The rest of the rules:
| Rule | What Kling says |
|---|---|
| Scope | Counted per account, model version and resource package type (video or image); all your API keys share it |
| When a slot is held | From the Submitted status until the task ends, failures included; released immediately after |
| Your cap | The highest concurrency among your active packages of the same type (a 5 and a 10 video package give 10, not 15) |
| Video or Virtual Try-on task | 1 slot each |
| Image task | As many slots as its n (n = 9 uses 9) |
| Requests per second | No QPS limit, per this page |
What is Kling API error 1303?
Code 1303 is the over-limit error: the number of running tasks has reached your concurrency cap. Kling's error-code table lists it under HTTP 429, and the concurrency page says it comes from system load, not from your parameters, so the request itself is fine to send again later.
The same table also lists 429 code 1302, "Too many requests; rate limit exceeded", and words 1303 as "Concurrency or QPS exceeds the prepaid resource package limit". That sits oddly with the concurrency page's "no QPS limits", so back off on both codes rather than assuming request rate is never refused. A 1303 body, from the concurrency page:
{
"code": 1303,
"message": "parallel task over resource pack limit",
"request_id": "..."
}How do I stay under the Kling concurrency limit?
Kling recommends two things, and its rules imply the rest:
- Retry 1303 with exponential backoff, starting at a delay of at least 1 second.
- Put a task queue in front of the create call, so you submit only as fast as slots free up.
- Size that queue to your package's concurrency, and count an image request as
nslots, not one. - Share one limiter across services: every API key on the account draws on the same quota.
- Poll as often as you need: status queries don't use a slot.
How do Kling API packages and units work?
You buy video and image generation resource packages, and Kling also offers a Trial Resource Package for integration testing. Usage is deducted from packages in units. One row of Kling's video price table as an example: Kling 3.0 without native audio at 1080P costs 0.8 Units ($0.112) per second. For the full per-second table, see Sume vs Kling.
To see what was deducted, POST /account/billing/package returns unit deductions per task, filterable by API key name, package name or ID, or product_type (video, image or try-on).
Can I run Kling without managing packages?
Through another API, yes, with that API's own limits. Sume's catalog lists kling-3, billed from the workspace's USD balance (Kling 3.0 API has the request and price). The behavior at the limit differs: when a Sume workspace is at its plan's concurrency, valid jobs are accepted as `queued` instead of refused, and only a full queue answers 429 queue_full.
Sources
Related posts
More in Developers
- Promise.allSettled vs Promise.all for a batch of API jobs
Promise.allSettled waits for every promise and reports each outcome; Promise.all rejects on the first failure. For paid API jobs, use allSettled.
- Python API rate limiting: stay under a per-minute limit
Pace Python API calls with an asyncio limiter set under the API's per-minute budget, keep polling on its own budget, and back off on 429 retry-after.
- Python requests default timeout: there isn't one
Python Requests has no default timeout: without timeout= a call can hang indefinitely. Set (connect, read) on every call, and keep it short for job APIs.
- Real time speech to text API: what a file-based API can do
Sume's speech to text API isn't real time: it transcribes recordings of up to 10 minutes at a URL. Chunked recordings give near-live transcripts.
Written by Sume