Replicate private models bill idle time: how Sume bills a job

On Replicate, a private model on dedicated hardware bills setup, idle and active time. Sume bills per job: a reserve at submit, then capture or refund.

4 min readSume
All posts

Replicate bills most private models for all the time their instances are online: the time spent setting up, the time spent idle waiting for requests, and the time spent processing. Only fast-booting fine-tunes are billed just for active time. Sume has no hardware you keep warm: a job reserves an estimated USD amount at submit, captures the actual cost on success, and refunds on failure or cancellation before capture.

Replicate's rules are from its Pricing page, read on 2026-10-02. Sume's are from the core workflow page, Video generation and Image generation.

How does Replicate bill public and private models?

Most public models are billed by the time they take to run, at a per-second price that depends on the hardware. Some are billed by input and output instead; the page's examples include black-forest-labs/flux-1.1-pro at $0.04 per output image and wavespeedai/wan-2.1-i2v-480p at $0.09 per second of output video.

Private models, made with Cog, mostly run on dedicated hardware, so you do not share a queue. The trade-off is that you pay for setup, idle and active time. Hardware is priced per second: the page lists gpu-h100 at $0.001525 per second, which it also shows as $5.49 per hour. Fast-booting fine-tunes, labeled in the model's version list, are billed only while active.

How does Sume bill a job?

Provider-backed generation reserves an estimated USD amount when you submit, captures the actual cost on success, and refunds on failure or cancellation before capture. The video docs describe a reserve at provider list price times 1.25 for every model, with usage.cost as the billable amount. The image docs say a generation is either completed and billed in full, or fails and is not billed, and that a client that disconnects early is billed as a failed generation, meaning not at all.

Sume's pages I read describe no hardware line, no idle charge and no setup charge; the unit is the job. That is a description of the docs, not a claim about every possible workload.

Which is the right shape for my workload?

The billing unit decides what you should measure.

Replicate's Pricing page and Sume's Core workflow, Video generation and Image generation docs, read 2026-10-02.
ItemReplicate private modelSume job
Billing unitHardware time: setup, idle, activePer job, by model rate
Idle timeBilled, except fast-booting fine-tunesNo idle line in the docs
FailureNot stated on the pageFailed generations are not billed; reserve refunded
Example rategpu-h100: $0.001525 per secondRead per model from the catalog (GET /v1/videos/models)
What to monitorUtilisation of the instanceSpend per job and wallet balance (/v1/usage)

What should I measure before choosing?

If your traffic is steady and you need a custom model, an always-on private instance has an hourly floor you can compute: hourly rate times hours online. If traffic is bursty or you only call catalog models, per-job billing has no floor, but every call has a price you should read first. The catalog's pricing fields (GET /v1/videos/models) show the rate before you spend.

Compute both numbers for your own month of traffic, not a headline rate.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume