What is serverless inference? How it works and is billed

Serverless inference means calling a hosted AI model over HTTP with no servers of your own. Providers bill for compute time or for each output.

5 min readSume
All posts

Serverless inference means you run an AI model by sending an HTTP request to a provider that hosts it, with no servers or GPUs of your own to provision, scale or keep running. The provider brings up hardware when requests arrive, and bills you either for the compute time your requests used or for each output the model returns.

The billing and cold-start facts below come from Replicate's and fal's own pages, read 2026-09-28 and listed under Sources. Sume's facts come from its video generation docs and generation admission docs.

What does inference mean here?

Inference is running a trained model on a new input: a prompt goes in, and an image, a video clip or a prediction comes out. Training builds the model; inference uses it. The serverless part only says who runs the machines. You never see an instance, only an endpoint, an API key and a bill.

An inference API is the HTTP interface to that model. Almost every serverless offering is one, so the two terms overlap. The difference that matters to a buyer is how the bill is counted.

How is serverless inference billed?

There are two billing shapes: compute time and per output. The same provider can use both, depending on whose model you run.

From Replicate: Billing, fal: Model API pricing and Sume's Video generation and Generation admission docs, read 2026-09-28.
Billing shapeWhat you pay forWhere the page says so
Compute time, public modelOnly the time the model is active on your requests; setup and idle time are freeReplicate, public models
Compute time, your own modelAll the time instances are online: setting up, idle and activeReplicate, most private models and deployments
Per outputEach output in the model's own unit (image, megapixel, video second); models with no fixed output price fall back to per-second GPU billingfal Model APIs
Per output, videoThe estimated amount at provider list × 1.25, reserved on submit and captured on successSume POST /v1/videos

Why does the billing shape matter?

On compute-time billing, a slow run costs more than a fast one, and on your own model an idle instance still costs money. On per-output billing, the unit is the thing you receive. fal says server errors (HTTP 500 or higher) are never billed and queue time is free (pricing). Sume reserves the estimate when it accepts a job, captures it on success, and releases or refunds it for failed jobs where applicable (admission). Sume video is billed from a workspace USD balance at provider list × 1.25, plus a 5.5% agent fee by default.

To compare two providers, price one real job on each: the same model, length and resolution. A headline rate in a different unit does not compare.

What is a cold start?

A cold start is the wait while a provider sets up an instance before it can run your request. Replicate describes it plainly: when requests arrive, it sets up an instance, which "can take a few seconds as we perform setup work like downloading weights". On a public model you share a hardware pool, so you "will sometimes encounter cold boots or scaling limits"; a deployment is its answer when you "want to avoid cold boots" (billing). fal's homepage says its serverless GPUs have "no cold starts" (fal).

Sume documents no cold-start behavior. What it documents is a queue: concurrency is a dispatch limit, not a submit limit, so a job beyond your plan's processing slots waits as queued until one opens. Video job concurrency and queueing has the per-plan numbers.

Can I run my own model serverless?

On some providers, yes. Replicate lets you "deploy your own custom models using Cog" (Replicate), and fal says billing for your own apps on its Serverless product "works differently" from its Model APIs (pricing); fal vs Replicate compares the two.

Sume is the catalog kind only. You pick a model id from GET /v1/videos/models, and "v1 runs a single backend per model", so you choose neither hardware nor provider, and its docs describe no way to deploy your own weights (Video generation).

Sources

Related posts

More in Developers

All Developers posts

Written by Sume