API uptime SLA: what it promises, and Sume's terms
An API uptime SLA promises a monthly availability percentage and credits if missed. Sume's terms offer none; how to build around that.

An API uptime SLA is a contract promise: the provider commits to a monthly availability percentage and owes you service credits if it misses it. Sume's public Terms make no uptime commitment. They provide the services "as is" and "as available", and say they may be unavailable, delayed, rate limited or degraded.
The Sume wording is quoted from its Terms of Service (last updated August 10, 2026). The retry advice comes from Errors and credits and Generation admission, all read 2026-09-29. This is what the pages say, not legal advice.
What does an uptime SLA actually promise?
Four things, and the percentage is only one of them. An SLA defines how uptime is measured, the percentage per month, what is excluded (for example planned maintenance), and the remedy, such as a credit against a future bill. It does not make outages impossible. It prices them.
The percentage turns into a monthly downtime budget:
| Monthly uptime | Downtime allowed |
|---|---|
| 99% | 7 h 12 min |
| 99.5% | 3 h 36 min |
| 99.9% | 43.2 min |
| 99.95% | 21.6 min |
| 99.99% | 4.3 min |
Does Sume have an API uptime SLA?
Not in its public Terms. What they say instead:
- The Services may be unavailable, delayed, rate limited, degraded, or changed from time to time.
- Agent runs and generation workflows depend on infrastructure and third-party services outside Sume's control.
- Sume does not guarantee that every request will be accepted, completed, or produce a usable result.
- The Services are provided "as is" and "as available", and Sume does not guarantee uninterrupted availability, error-free operation or provider availability.
Can an enterprise contract include one?
Sume's Terms say Enterprise services may be governed by a signed Order Form, and that where its commercial terms conflict with the Terms, the signed agreement controls. No public page says what availability terms an Order Form offers, so ask in writing. AI video generator for enterprise covers the rest of that process.
How do I build on an API without an SLA?
Assume individual requests will sometimes be refused, and make every refusal safe to retry. For Sume's generation API:
- Send an
Idempotency-Keyon every create. On a Format call, same key and same body return the original receipt: no second run, no second charge. - On
429, back off and useretry-afterwhen present.429 queue_fullmeans your workspace's concurrency plus queue is full, not that the service is down. - On
503 provider_capacity_exceeded, retry later. The docs say to reuse the same key, but in current code a same-key retry afterqueue_fullorprovider_capacity_exceededreplays the refusal. Use a new key once the cause clears. - On
provider_not_configured, don't retry aggressively. - Keep the Sume request id from the response body or headers and include it when you report an issue.
What does a 429 or 503 tell me about uptime?
A 429 is about your limits, a 503 is about Sume's side. Neither is an outage by itself, and a submission accepted as queued is working as designed: Sume accepts valid jobs as queued while queue capacity remains. 429 vs 503: rate limit or server overload? explains the difference in detail.
Sources
Related posts
More in Developers
- Automate video editing in Python with an editing API
Automate video editing in Python by calling an editing API with Requests: submit a caption, cut, or crop job, poll until it ends, chain the output.
- Bash for loop with curl: one API request per line
Loop over a file with while IFS= read -r, build each JSON body with jq --arg, send it with curl --fail-with-body, and pace it under the API's rate limit.
- Batch video processing by API: one edit, many videos
Batch video processing by API is a loop: one edit job per file, each with its own idempotency key, collected by webhook. How it works on Sume, and costs.
- C# HttpClient default timeout: 100 seconds, and how to set it
HttpClient.Timeout defaults to 100 seconds per request and throws TaskCanceledException. How to set it, and why slow API jobs need polling instead.
Written by Sume