Replicate official models: always warm, stable API. What Sume promises
Replicate says official models are always warm with a stable input and output API, priced by output. Sume's catalog ids, canonical_slug and retirement notices.

Replicate's official models page makes three promises: the input and output API of every official model is stable, official models are always warm so there are no cold boots, and they are billed by output (per image, token or similar) rather than by runtime. Sume offers a catalog with permanent identifiers and per-output pricing, but its docs do not promise "always warm", and they openly list routes that are retiring.
This post lines the claims up, using only what each vendor's page says.
What does Replicate say about official models?
From the page read on 2026-10-02: official models are "charged by output (or in some cases, input)" with the metric on each model page; "the input and output API for every official model is stable"; they are "always warm and ready to respond to requests, so you can run them without worrying about cold boots"; and Replicate maintains over 100 of them. In client libraries you call an official model by owner/name with no version, and over HTTP the route is POST /models/<owner>/<name>/predictions.
Replicate's pricing page separately says most public models bill by hardware and duration, and some by output, and that private models typically pay for setup and idle time as well. Official models are the output-priced, warm option.
What does Sume document?
The contrast is not like for like, because the Sume docs describe a catalog rather than a community marketplace. What they document is a catalog read from the API: GET /v1/videos/models and GET /v1/images/models return each model's id, a canonical_slug described as the permanent model identifier, supported resolutions, aspect ratios and durations, and pricing_skus. Billing is a reservation on submit at provider list price times 1.25, and usage.cost is the Sume billable amount.
Two things Sume does not claim: it does not promise warm capacity or latency, and it says outright that limits vary by model (for example seedance-2.5 takes 4 to 30 seconds, most others stop at 15). Read capabilities from the catalog, not from memory.
| Question | Replicate official models | Sume |
|---|---|---|
| Billing basis | Per output, shown on each model page | Per output second or image, list x 1.25, reserved on submit |
| Warm capacity | "Always warm" | Not promised; jobs queue under plan concurrency |
| Input and output stability | Stated stable for every official model | Per-model capabilities in the catalog; compatibility aliases kept |
| Version in the call | None for official models | Bare catalog id such as seedance-2 |
| Retirements | Not covered on this page | Documented, for example Video 1.0 and Music 1.0 |
What retires on Sume?
Honesty about change is part of stability. Sume's docs mark Video 1.0 (POST /v1/video-1.0/generate) and Image 1.0 as retiring soon, kept as compatibility aliases for the Auto pipeline, and Music 1.0 as retiring and resolving through Music Router. The Video Router generate route stays available and unchanged, but new integrations are told to use POST /v1/videos, which is the same catalog and jobs behind an OpenRouter-compatible wire.
So the stability you can lean on at Sume is the catalog contract and the compatibility aliases, not a promise that every route lives forever.
How does "no cold boots" show up in a client?
On Replicate it means a prediction on an official model should start without waiting for a boot. On Sume, the equivalent question is whether your job sits in queued. A job waits there when your workspace is at its plan's processing concurrency (Free 1, Pro 4, Startup 8, Scale 20), and the docs say not to treat queued as a failure. Poll with backoff, or use a webhook, and never resubmit just because a local timer expired.
curl https://api.sume.com/v1/videos/models \
-H "Authorization: Bearer $SUME_API_KEY"
# read: id, canonical_slug, supported_durations, pricing_skus
curl https://api.sume.com/v1/jobs/job_123/status \
-H "Authorization: Bearer $SUME_API_KEY"Which should you pick?
If you want a large set of maintained models with a stable per-model API and output pricing, Replicate's official models are a direct fit. If you want one wire shape across image, video and music with an Auto option and a margin-inclusive price visible before you submit, Sume is the alternative. Neither page gives benchmark numbers, so test both on your own prompts. The broader comparison is in Sume vs Replicate; Sume's catalog details are in Video generation.
Sources
Related posts
More in Comparisons
- Replicate predictions time out at 30 minutes: Sume's deadline
Replicate stops a prediction after 30 minutes unless support raises it. Sume documents no such field: the deadline is client-side and a timeout does not cancel.
- Replicate status succeeded vs Sume completed: map the states
Replicate uses starting, processing, succeeded, failed, canceled. Sume uses queued, processing, completed, failed, canceled. A mapping table and a poll loop.
- Replicate Prefer: wait holds 60 s by default; Sume's sync cap is 30 s
Replicate's Prefer: wait header holds the request up to 60 seconds by default; Sume's sync mode waits at most 30 seconds, then returns a job to poll.
- Replicate SSE stream URL vs Sume: no event stream, so poll
Replicate returns a urls.stream endpoint with output, error and done events. Sume has no SSE or WebSocket on the Developer API. Poll status or take a webhook.
Written by Sume