Check price before submit: fal pricing endpoint vs Sume video models

fal prices by endpoint_id at GET /v1/models/pricing. Sume lists pricing_skus in GET /v1/videos/models, for example 0.0154 per 1000 video tokens on seedance-2.

4 min readSume
All posts

Both platforms let you look up a price before you submit. fal exposes GET https://api.fal.ai/v1/models/pricing?endpoint_id=.... Sume's GET /v1/videos/models rows carry a pricing_skus object, and pricing lines on each image endpoint record. Neither replaces a dry run on your actual request, but both let a script refuse a job that is too expensive before it queues.

What each side documents

Pricing lookups, read 2026-10-05
ItemfalSume
LookupGET /v1/models/pricing?endpoint_id=...GET /v1/videos/models; GET /v1/images/models/{id}/endpoints
Billing basisPrepaid credits; successful outputs onlyWallet; images all-or-nothing, failed or canceled images are not billed
Exampleflux/dev $0.025 per imageseedance-2: per-1000-video-tokens 0.0154 (docs example response)
UnitPer image for the example abovePer 1,000 video tokens; image lines are cost_usd per image

Arithmetic

fal: 40 images at $0.025 each is 40 x 0.025 = $1.00. Sume image: the docs example endpoint line is cost_usd 0.033 per output image, and you pay cost_usd x n, so n = 4 is 4 x 0.033 = 0.132. Sume video: the catalog value is 0.0154 for each 1,000 video tokens. For an illustrative job of 100,000 tokens (an assumed number to show the math, not a docs figure), 100 x 0.0154 = 1.54. The docs example does not state the token count for any clip length, so read the live catalog and do a dry run before relying on a total.

The docs example row is a sample response, so the live values can differ. Image endpoint pricing lines already include the Sume margin.

A Python check against the live catalog

This reads the video catalog and prints each model's SKUs. A budget guard can compare them before you submit.

import asyncio, os
import httpx

async def main():
    headers = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
    async with httpx.AsyncClient(timeout=30) as client:
        r = await client.get("https://api.sume.com/v1/videos/models", headers=headers)
        r.raise_for_status()
        for m in r.json()["data"]:
            print(m["id"], m.get("pricing_skus"))

asyncio.run(main())

Where to be careful

  • A SKU is a rate, not a quote. The job size decides the total.
  • fal bills successful outputs only and keeps results about an hour; Sume releases or refunds the reservation on failed jobs. Both follow the pay-for-success idea, but check each page for the exact rule.
  • For a hard cap on calls to Sume's hosted MCP, add max_spend_usd or use dry_run.

Building a budget guard

A guard needs three inputs: the rate (from the pricing lookup), your estimate of the job size, and a ceiling. Multiply the first two and compare with the third. For fal, the rate is per endpoint; for Sume, read pricing_skus and, for images, the pricing lines. When the estimate is hard to make, such as tokens for a video, use the Sume dry_run or generation_admission_preview over MCP, which return an estimate without creating a job.

Log the rate and the estimate with the job id, so a surprise on the invoice can be matched to a decision your code made.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume