fal H100 at $2.49/hour vs per-output media API pricing: which fits

fal lists an H100 at $2.49 per hour and video APIs from $0.025 to $0.14 per second. When hourly GPU pricing beats paying per output, and when it does not.

5 min readSume
All posts

fal's pricing page lists serverless GPUs by the hour (an H100 at $2.49 per hour, list $4.50) and, separately, model APIs billed per output, with video from $0.025 to $0.14 per second. Hourly pricing wins only when you run your own model and keep the GPU busy; per-output pricing wins when you want a result and not a machine, which is the model Sume uses.

What does fal's page list for GPUs?

The page shows GPU rates with a discounted price and a list price. These are the figures it showed on 2026-10-02.

fal serverless GPU prices per hour, read 2026-10-02
GPUMemoryPrice per hourList per hour
B300288GB$5.99$12.99
GB200192GB$5.89$9.99
B200192GB$5.49$7.99
H200141GB$2.99$6.00
H10080GB$2.49$4.50
RTX PRO 600096GB$1.99$4.00

What is the break-even between hourly and per-output?

An H100 at $2.49 per hour is about $0.0007 per second of wall-clock time ($2.49 divided by 3,600). If your own model needs 90 seconds of GPU time to make a clip, that is roughly $0.06 of compute, but only when the GPU is fully busy. Cold starts, idle time between jobs, failed runs and the engineering to host the model all land on you.

The page's own model API examples show the other side: video from $0.025 to $0.14 per second, and the page notes that "settings such as resolution, duration, and quality may affect the final cost". A five-second clip at the top of that range is $0.70; at the bottom it is $0.125. Hourly pricing is the better deal if you are generating all day on a model you control. It is the worse deal for bursty work.

How does Sume price a job?

Sume does not sell GPU hours. You submit a job and Sume reserves the estimated USD cost from your wallet, then captures it on completion or releases it on failure, per Generation admission. For listed models the rule is provider list price times 1.25.

That means a bursty month costs what you generated and nothing for idle time. Sume's docs describe a catalog of listed models, not GPU hours or custom weights, so a fine-tuned model of your own needs a different home.

curl https://api.sume.com/v1/balance \
  -H "Authorization: Bearer $SUME_API_KEY"

curl "https://api.sume.com/v1/usage?limit=20" \
  -H "Authorization: Bearer $SUME_API_KEY"
# reserved, captured and refunded rows show what a job held and what it kept

Which should you pick?

Pick hourly GPUs when you have a custom or fine-tuned model, steady volume and someone to operate it. Pick per-output when you want catalog models, predictable per-job cost and no idle bill. Either way, compare on cost per kept result, not the hourly or per-second sticker: divide spend by the clips you actually ship.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume