Replicate private models bill idle time: how Sume bills a job
On Replicate, a private model on dedicated hardware bills setup, idle and active time. Sume bills per job: a reserve at submit, then capture or refund.

Replicate bills most private models for all the time their instances are online: the time spent setting up, the time spent idle waiting for requests, and the time spent processing. Only fast-booting fine-tunes are billed just for active time. Sume has no hardware you keep warm: a job reserves an estimated USD amount at submit, captures the actual cost on success, and refunds on failure or cancellation before capture.
Replicate's rules are from its Pricing page, read on 2026-10-02. Sume's are from the core workflow page, Video generation and Image generation.
How does Replicate bill public and private models?
Most public models are billed by the time they take to run, at a per-second price that depends on the hardware. Some are billed by input and output instead; the page's examples include black-forest-labs/flux-1.1-pro at $0.04 per output image and wavespeedai/wan-2.1-i2v-480p at $0.09 per second of output video.
Private models, made with Cog, mostly run on dedicated hardware, so you do not share a queue. The trade-off is that you pay for setup, idle and active time. Hardware is priced per second: the page lists gpu-h100 at $0.001525 per second, which it also shows as $5.49 per hour. Fast-booting fine-tunes, labeled in the model's version list, are billed only while active.
How does Sume bill a job?
Provider-backed generation reserves an estimated USD amount when you submit, captures the actual cost on success, and refunds on failure or cancellation before capture. The video docs describe a reserve at provider list price times 1.25 for every model, with usage.cost as the billable amount. The image docs say a generation is either completed and billed in full, or fails and is not billed, and that a client that disconnects early is billed as a failed generation, meaning not at all.
Sume's pages I read describe no hardware line, no idle charge and no setup charge; the unit is the job. That is a description of the docs, not a claim about every possible workload.
Which is the right shape for my workload?
The billing unit decides what you should measure.
| Item | Replicate private model | Sume job |
|---|---|---|
| Billing unit | Hardware time: setup, idle, active | Per job, by model rate |
| Idle time | Billed, except fast-booting fine-tunes | No idle line in the docs |
| Failure | Not stated on the page | Failed generations are not billed; reserve refunded |
| Example rate | gpu-h100: $0.001525 per second | Read per model from the catalog (GET /v1/videos/models) |
| What to monitor | Utilisation of the instance | Spend per job and wallet balance (/v1/usage) |
What should I measure before choosing?
If your traffic is steady and you need a custom model, an always-on private instance has an hourly floor you can compute: hourly rate times hours online. If traffic is bursty or you only call catalog models, per-job billing has no floor, but every call has a price you should read first. The catalog's pricing fields (GET /v1/videos/models) show the rate before you spend.
Compute both numbers for your own month of traffic, not a headline rate.
Sources
Related posts
More in Comparisons
- Replicate's six prediction statuses vs Sume's five job statuses
Replicate has starting, processing, succeeded, failed, canceled and aborted. Sume has queued, processing, completed, failed and canceled. The map and its gap.
- Resemble AI vs Sume: voice, detection and watermarking vs video
Resemble AI covers TTS, speech-to-speech, deepfake detection and watermarking. Sume makes media but ships no detection API. What each covers.
- Respeecher Space at $2 an hour vs Sume async text to speech
Respeecher Space is a real-time TTS API for voice agents at $2 an hour. Sume TTS is async and per character. Which fits narration, and which fits live voice.
- Runway lists 18 video model ids: which nine does Sume have?
Of the 18 video ids on Runway's models page, nine match a Sume id: four Seedance rows, MiniMax H3 and Max, Wan 3.0, Grok Imagine, Gemini Omni Flash 1.1.
Written by Sume