fal vs Replicate: how calls, webhooks and billing differ

fal and Replicate both run many models behind one async API. fal bills per output and recommends its queue; Replicate bills most models by time.

5 min readSume
All posts

fal and Replicate both run many third-party models behind one HTTP API, with async jobs, webhooks and a blocking option. The biggest difference is billing: fal bills per output (video per second or per video) from prepaid credits, while Replicate bills most models by hardware time per second and some by input and output.

Every fal and Replicate fact below comes from that vendor's own docs or pricing page, read on 2026-09-28 and listed under Sources. Neither is ranked; which fits depends on how you call models and how you want to be billed. If you are looking past both, fal.ai and Replicate alternatives for AI video covers that question.

How do fal and Replicate differ?

Both are model hosts, not model makers: you pick a model by id and send its inputs. The differences are in the plumbing around the call.

From fal's queue, synchronous, webhooks, pricing and client libraries docs, and Replicate's prediction, webhooks, billing and client libraries docs, read 2026-09-28.
ItemfalReplicate
Catalog1,000+ image, video, audio and 3D modelsThousands of community models, plus official models
Default callQueue (recommended): poll or take a webhook; run calls fal.run directly and subscribe blocks while it polls the queueAsync prediction by default; Prefer: wait holds the request open, 60 seconds by default
Webhookswebhook_url; a POST with the result when the request completes, signed with ED25519webhook on create; POSTs when the prediction is created, updated and finished
Billing unitPer output: video per second or per video; GPU seconds for models with no output priceMost models by hardware time per second; some by input and output
FailuresHTTP 500+ errors and queue wait are not billedFailed runs are not charged; a canceled official model may still be
PaymentPrepaid creditsPrepaid credit upfront; some accounts billed in arrears
Client librariesPython, JavaScript, Swift, Kotlin / Java, DartNode.js, Python, Swift, Go, plus an MCP client
Your own modelsfal ServerlessCog, an open-source packaging tool

How does each one wait for a result?

fal calls asynchronous inference “the recommended way to call models on fal”. For simple blocking calls it also offers run, a direct request to fal.run with no queue, and subscribe, which polls the queue for you (synchronous inference). A queue submit returns a request_id plus a status_url. A request moves through IN_QUEUE, IN_PROGRESS and COMPLETED. fal says queued requests are never dropped, and a request whose runner fails is re-queued and retried up to 10 times (queue). Webhook deliveries are retried up to 31 times until the stored result expires: about 1 hour after the request completes, or about 6 minutes for results of 10 KB or more (webhooks).

Replicate creates predictions in async mode unless you ask otherwise. With the Prefer: wait header it holds the request open, 60 seconds by default, and returns the unfinished prediction if the model runs longer, so you poll it after all (create a prediction). A Cancel-After header sets a deadline between 5 seconds and 24 hours. Input and output files of API predictions are deleted after an hour, so save what you keep (webhooks).

How does fal vs Replicate pricing work?

The two bill different things, so compare a real job, not a headline rate.

  • fal: “you are billed based on the output you generate.” Video models charge per second of generated video or a flat rate per video, and models with no fixed output price fall back to per-second billing by GPU type (pricing).
  • Replicate: “Most models are billed by the time they take to run,” at a price per second set by the hardware; some models are billed by input and output instead (pricing). On a public model you pay only while it is processing your requests (billing).
  • Both sell prepaid credit. fal credits also fund your concurrency limit (pricing); Replicate asks you to “purchase credit upfront” (prepaid credit), and bills some accounts in arrears instead (billing).

Can I run my own model on either?

Yes. fal lists fal Serverless for custom models and on-demand GPUs next to its model APIs (fal). Replicate lets you “deploy your own custom models using Cog,” its open-source packaging tool (Replicate). Your own deployments are billed differently from the public catalog on both: fal says Serverless billing “works differently” (pricing), and Replicate bills private models for all the time their instances are online, setup and idle included (billing).

Where does Sume fit if I only need video?

Sume is a third option with a narrower catalog: POST /v1/videos returns a job id and a polling_url, takes an optional callback_url for a webhook, and reserves each job at the provider's list price × 1.25 from one workspace balance (Video generation). Sume vs fal and Sume vs Replicate compare each one with Sume.

Which should I pick?

  • Pick fal if you want per-output prices, a queue with automatic retries, and clients for mobile platforms (Swift, Kotlin / Java, Dart).
  • Pick Replicate if billing by run time suits your models, you want a Go client, or you want to package your own model with Cog.
  • Pick either for the same popular model: check that model's page on each, since the unit (per second of video, per output, or per hardware second) decides the bill.
  • Pick Sume's video endpoint if the video models it lists are all you need.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume