Replicate predictions time out at 30 minutes: Sume's deadline
Replicate stops a prediction after 30 minutes unless support raises it. Sume documents no such field: the deadline is client-side and a timeout does not cancel.

Replicate times a prediction out after 30 minutes of running, and asks you to contact support for longer; Sume's docs describe no server-side run deadline you can set. On Sume the overall deadline is your client's, and a client-side timeout does not cancel the job, which keeps running and billing.
Replicate's rule is on its prediction lifecycle page; Sume's are in Jobs and results and Generation admission, all read on 2026-10-02.
What does the 30-minute limit mean on Replicate?
The lifecycle page says predictions time out after running for 30 minutes and that longer timeouts need a support request. You can also set your own deadline so a prediction stops earlier. The billing rule depends on where the stop lands: a prediction aborted before it started costs nothing, while one canceled after starting is billed for the runtime it used.
That makes the 30 minutes a cost ceiling per prediction as much as a time limit.
Does Sume have an equivalent?
The pages we read do not list a per-job run limit or a deadline field. They do say how long each wait can be: a sync submit holds at most 30 seconds (wait_timeout_seconds is clamped to 0 to 30), and hosted MCP jobs_wait calls hold at most 55 seconds. Anything longer is polling or a webhook.
The docs call 20 minutes a reasonable client-side deadline for video, and state plainly that it is not wait_timeout_seconds.
| Question | Replicate | Sume |
|---|---|---|
| Run limit | 30 minutes, raise via support | None documented |
| Who sets a shorter one | Prediction deadline you set | Your client |
| Longest HTTP hold | Not covered here | 30 s sync, 55 s jobs_wait |
| After a client timeout | Prediction continues until its limit | Job keeps running and billing |
How do you enforce a deadline on Sume?
Poll against your own clock, and cancel explicitly if you give up. Cancel only works before generation starts; once it has started, Sume returns 409 job_generation_already_started and the job completes or fails normally. So a deadline you want to enforce must be paired with a cancel while the job is still queued.
This sketch stops watching at your deadline; add a POST /v1/jobs/:id/cancel call in the handler if the job may still be queued.
import os, time, requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
BASE = "https://api.sume.com/v1/jobs"
def wait(job_id, deadline_s=1200):
end, delay = time.time() + deadline_s, 5
while time.time() < end:
r = requests.get(f"{BASE}/{job_id}/status", headers=H, timeout=30)
r.raise_for_status()
s = r.json()
if s.get("terminal"):
return s
time.sleep(s.get("next_poll_after_seconds") or delay)
delay = min(delay * 2, 30)
raise TimeoutError(f"still running: {job_id}")What does this change in a port?
Drop code that expects the platform to kill a long run. Add a budget instead: Sume reserves its USD estimate at submit time and captures it on success, so the spend you can lose to a slow job is bounded by the estimate, not by minutes. See the Replicate cancel-after comparison for the header-level version.
Sources
Related posts
More in Comparisons
- Replicate private models bill idle time: how Sume bills a job
On Replicate, a private model on dedicated hardware bills setup, idle and active time. Sume bills per job: a reserve at submit, then capture or refund.
- Replicate's six prediction statuses vs Sume's five job statuses
Replicate has starting, processing, succeeded, failed, canceled and aborted. Sume has queued, processing, completed, failed and canceled. The map and its gap.
- Resemble AI vs Sume: voice, detection and watermarking vs video
Resemble AI covers TTS, speech-to-speech, deepfake detection and watermarking. Sume makes media but ships no detection API. What each covers.
- Respeecher Space at $2 an hour vs Sume async text to speech
Respeecher Space is a real-time TTS API for voice agents at $2 an hour. Sume TTS is async and per character. Which fits narration, and which fits live voice.
Written by Sume