Replicate deletes API outputs after an hour: what Sume keeps
Replicate removes API prediction inputs, outputs and files after an hour by default. Sume result URLs on media.sume.com do not expire. What to save and when.

How long does Replicate keep API prediction output?
One hour by default. Replicate's data retention page says that for predictions created through the API, all input parameters, output values, output files and logs are automatically removed after an hour. Predictions made in the web interface are kept indefinitely, and you can delete any prediction manually from the dashboard.
Sume works the other way for generated media: completed jobs return public media.sume.com artifact URLs, and Sume's run documentation says those URLs do not expire. Both are design choices, and each one decides where your pipeline has to be careful.
What does the one-hour window change in your code?
Replicate's prediction object includes a data_removed field, and its reference says that after removal the output key persists but returns null. Its webhooks page frames webhooks as the way to persist data before automatic deletion, since a webhook delivers the metadata for a completed prediction so you can save output files to your own storage.
The practical rule on Replicate: treat the output URL as a transfer handle, not a storage location. Download or copy the file as soon as the prediction succeeds, and poll with the clock in mind, because a slow consumer that returns after an hour finds nothing.
What does Sume return, and what is the catch?
A completed Sume job exposes artifacts with url, type and content_type, as in Jobs and results. The run receipt docs state that media URLs are durable media.sume.com HTTPS URLs that do not expire and are public to anyone holding the URL. See Run receipts.
The catch is the other half of that sentence. A durable, public URL is not an access-controlled one. If your product needs per-customer access, proxy or copy the file into your own storage and serve it from there. Sume does not document a private or signed-URL mode for these artifact URLs in the pages above, so do not assume one.
Comparison at a glance
Retention facts are from each vendor's own pages.
| Question | Replicate (API predictions) | Sume (job artifacts) |
|---|---|---|
| Default retention | Removed after an hour | Media URLs do not expire |
| Web UI results | Kept indefinitely | Not applicable to the API |
| What disappears | Inputs, outputs, files and logs | Nothing on a timer |
| Who can fetch the URL | Not stated on the pages read | Anyone who holds the URL |
| Your obligation | Copy files out within the hour | Copy only if you need access control |
Which should you plan around?
If you cache outputs in your own bucket anyway, the difference is small: save on completion and move on. The difference shows up in three places.
- Retries and slow consumers: on Replicate a queue backlog can outlive the output; on Sume a late reader still gets the file.
- Audit trails: a Sume job id keeps resolving to its result and its usage rows; a Replicate record loses its content after the window.
- Privacy: short retention is a feature if you do not want vendor-side copies. Sume says nothing on that page about deleting a URL on request, so ask support before promising a customer erasure.
Sources
Related posts
More in Comparisons
- Replicate Cancel-After header: 5 s to 24 h. Is there a Sume deadline?
Replicate's Cancel-After header sets a prediction deadline from 5 seconds to 24 hours. I found no deadline field in Sume's OpenAPI, so cancel it yourself.
- Replicate predictions time out at 30 minutes: Sume's deadline
Replicate stops a prediction after 30 minutes unless support raises it. Sume documents no such field: the deadline is client-side and a timeout does not cancel.
- Replicate status succeeded vs Sume completed: map the states
Replicate uses starting, processing, succeeded, failed, canceled. Sume uses queued, processing, completed, failed, canceled. A mapping table and a poll loop.
- Replicate Prefer: wait holds 60 s by default; Sume's sync cap is 30 s
Replicate's Prefer: wait header holds the request up to 60 seconds by default; Sume's sync mode waits at most 30 seconds, then returns a job to poll.
Written by Sume