fal's Usage API vs Sume's per-job, per-run cost reads

fal's changelog lists a Usage API for cost attribution. Sume's GET /v1/usage filters by job_id, run_id or thread_id and returns a summary. How to attribute.

4 min readSume
All posts

To attribute spend on Sume, call GET /v1/usage with job_id, run_id or thread_id and quote the summary.debited_usd figure, instead of adding up ledger rows yourself. The fal changelog lists a Usage API for cost attribution alongside per-condition retry budgets, termination grace up to one hour and a Platform MCP Server for debugging failed requests.

The two services solve the same bookkeeping problem, and the Sume docs are firm on one point: never sum rows yourself.

What the fal page lists

Only the items relevant to cost and failures are used.

fal changelog (read 2026-10-03)
ItemWhat the page lists
RetriesPer-condition retry budgets
TerminationGrace period up to 1 hour
DebuggingPlatform MCP Server for failed requests
CostUsage API for cost attribution

What the Sume summary contains

When you scope GET /v1/usage to one thread, run or job, the response adds a summary folded over every ledger row the scope caused. limit only caps the rows listed, not the summary.

Sume usage summary fields, from the usage docs (read 2026-10-03)
FieldMeaning
debited_usd, debited_usd_microsWhat the wallet deducted; the figure to quote
held_usd_microsHolds still open; not spend yet
refunded_usd_microsHolds given back after failure, cancellation or queue_full; not spend
finaltrue once no hold is open
by_operation_type, runsThe same money by operation type, and by run for a thread

Why not sum the rows

A refunded row keeps its hold amount in billable_amount_usd_micros, so a naive sum counts money you got back. The summary already excludes it. Rows also carry thread_id, run_id, turn_job_id and script_run_id, so you can group them yourself for reporting, but the quoted total should come from the summary.

job_id also accepts a turn's job id, which sums the turn's own row plus every job it commissioned.

A per-job lookup

The request below reads the cost of one job. Replace the id with one you stored at submit time. The live OpenAPI schema is the source of truth for exact fields.

curl "https://api.sume.com/v1/usage?job_id=job_123" \
  -H "Authorization: Bearer $SUME_API_KEY"

Attribution that survives retries

The reservation model makes the retry story simple.

  • A reservation is held when a job is accepted and captured on success.
  • A failed job, or a failed queue admission, releases the reservation where applicable.
  • A retry with the same Idempotency-Key returns the original job, so there is one ledger story per intent.
  • Store your own customer or campaign id next to the Sume job id at submit time; Sume's rows do not know your tags.

Retries and failures

fal documents retry budgets and a termination grace period on its side. On Sume the equivalents are narrower: provider_capacity_exceeded is retried later with the same idempotency key, queue_full calls for backing off, a generation_timeout or worker_timeout should lead you to poll status, and cancellation succeeds only before generation starts.

In both services the cost question after a failure is the same: was money captured or released. On Sume the final flag in the summary tells you when no hold is open.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume