Replicate predictions by model type vs Sume's one route per product

Replicate creates predictions through different calls for community models, official models and deployments. Sume uses one route per product. What to map.

5 min readSume
All posts

On Replicate the call you make depends on what you are running: predictions.create for community models, models.predictions.create for official models and deployments.predictions.create for deployments. On Sume you call one route per product, such as POST /v1/videos or POST /v1/images, and choose the model with a model field in the body.

That is the whole structural difference, and it decides how much routing code your client needs.

What are Replicate's three ways to create a prediction?

The Replicate create-a-prediction page names three client methods by model type: community models use predictions.create, official models use models.predictions.create, and deployments use deployments.predictions.create. The official-models page adds that over HTTP the official route is POST /models/<owner>/<name>/predictions and that you call official models without a version.

The same page describes two modes. Sync mode waits up to 60 seconds by default, adjustable with a Prefer: wait=X header, and returns the output if the prediction finishes in time. Async mode returns immediately with a prediction id. A Cancel-After header sets a deadline between 5 seconds and 24 hours, and webhooks accept a webhook_events_filter such as ["completed"].

What does the same step look like on Sume?

Sume's endpoint map is organized by product, not by who hosts the model. Video generation is POST /v1/videos with a model such as seedance-2.5 or sume/auto; the older Video Router is POST /v1/video-router/generate; images are POST /v1/images; music is POST /v1/music-router/generate. Every one returns the same job envelope, and the same job is readable at GET /v1/jobs/{id}/status and /result.

Modes are async (the default), webhook, and sync or subscribe, which wait at most 30 seconds. There is no deployment concept in the public API: you do not create or scale a private model.

Creating a generation, Replicate versus Sume (read 2026-10-02)
StepReplicateSume
Pick a modelRoute or method differs by model typemodel field in the body
Run asyncDefault; returns a prediction idmode: "async" (default); returns a job id
Wait inlinePrefer: wait, 60 seconds by defaultsync up to 30 seconds, then poll
DeadlineCancel-After, 5 seconds to 24 hoursNo deadline field; cancel before start
Private capacityDeploymentsNot offered; plan concurrency applies

How do I translate a Replicate client?

Collapse the three branches into one function that takes a model id and a body. Everything Replicate expressed through the route becomes data on Sume. A deployment name has no equivalent, so decide which catalog model it was standing in for.

Replace Prefer: wait with a poll loop keyed on terminal and treat a sync timeout as "keep polling" since the docs say a non-terminal answer means poll, never resubmit. Replace Cancel-After with your own client-side deadline plus a POST /v1/jobs/{id}/cancel call, remembering that cancel succeeds only before generation starts and otherwise returns 409 job_generation_already_started.

import asyncio, os, httpx

H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

async def run(model: str, prompt: str) -> dict:
    async with httpx.AsyncClient(base_url="https://api.sume.com", headers=H) as c:
        r = await c.post("/v1/videos", json={"model": model, "prompt": prompt},
                         headers={"Idempotency-Key": "demo-0001"})
        job = r.json()
        while job.get("status") not in ("completed", "failed", "cancelled"):
            await asyncio.sleep(10)
            job = (await c.get(job["polling_url"])).json()
        return job

async def main():
    print(await run("sume/auto", "A product clip on a desk"))

asyncio.run(main())

What is lost and gained?

You lose the community long tail, deployments and the per-request deadline header. You gain a single request shape, an idempotency key that makes retries safe (a replay returns the original job), and a price shown before you submit at provider list times 1.25. Neither side offers a streaming progress channel for video on these pages that this post relies on; Sume's docs say there is no SSE stream and progress comes from polling events_url.

For the rest, read Sume vs Replicate and Sume models and routes.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume