Which Sume API calls are free: balance, usage, catalog, filter check
Balance, usage and catalog reads cost nothing on Sume, and neither does the video-filter check. What is billed, what is refunded, and what a 402 means.

On Sume, the calls that never spend your balance are the read-only ones: GET /v1/balance, GET /v1/usage, GET /v1/catalog, and the program check at POST /v1/video-filter/check, which the docs call unbilled and which creates no job and reserves no credits. Money only moves when a paid generation job is accepted: Sume reserves the estimate at submit, captures it on success and releases it on failure or an early cancel.
That answers the usual worry about polling. Status checks are rate limited, not priced. The docs treat read and status endpoints as polling backpressure, and give reads a budget 40 times the write budget so that polling does not block your submits. The rest of this post lists what is free, what is billed, and when you get money back.
Which calls are free?
The Usage docs describe GET /v1/balance and GET /v1/usage as read-only and scoped to the workspace behind the API key. The catalog read returns models and their pricing metadata. The video-filter check runs the same schema and source preflight as the real encode, returns diagnostics and, when the program is valid, an estimate, and the docs say it does not create a job, reserve credits, boot a box or touch the encoder.
One caution: free does not mean unlimited. Every authenticated call counts against a per-key rate budget, and a 429 rate_limited response carries retry-after.
| Call | Moves money? | Notes |
|---|---|---|
GET /v1/balance | No | USD balance in micros and cents; state is funded or empty |
GET /v1/usage | No | Ledger rows; job_id, run_id or thread_id filters add a summary |
GET /v1/catalog | No | Includes estimated, minimum and maximum cents per model |
POST /v1/video-filter/check | No | No job, no reservation; returns an estimate when valid |
GET /v1/jobs/:id/status | No | Read; rate limited, not priced |
POST /v1/video-filter | Yes | Flat $0.02 per encode job, reserved at submit |
| Any paid generation submit | Yes | Reserved at submit, captured on completion |
When does a billed call give the money back?
The Usage docs list three ledger statuses: reserved for an estimate held before provider work, captured for billable usage taken after success, and refunded for a reservation released after a failure or a cancel before capture. A job that fails does not keep its hold. A 429 queue_full releases the reservation made for that failed attempt. Cancelling works only before generation starts; after that the call returns 409 job_generation_already_started and the job completes or fails normally.
A 402 insufficient_credits means the balance could not cover the estimate for your request. The OpenAPI text is explicit that no generation job is started, so there is nothing to refund and nothing to cancel.
- Failed job: reservation released.
- Cancel before generation starts: reservation released.
queue_fullon submit: reservation released for that attempt.402: refused before provider work; nothing held.- Idempotent retry with the same key and body: the same job, not a second charge.
How do I see exactly what one job cost?
Add job_id, run_id or thread_id to GET /v1/usage and the response gains a summary. The field to quote is debited_usd_micros, what the wallet actually deducted. held_usd_micros is holds still open, not yet spend, and refunded_usd_micros is money given back. The final flag is true once no hold is open. The docs warn against summing rows yourself, because a refunded row keeps its hold amount in billable_amount_usd_micros.
curl https://api.sume.com/v1/balance \
-H "Authorization: Bearer $SUME_API_KEY"
curl "https://api.sume.com/v1/usage?job_id=$JOB_ID&limit=50" \
-H "Authorization: Bearer $SUME_API_KEY"What can I use the free calls for?
Use them to avoid paid mistakes. Call the filter check before a filter job, since a program that passes the check can still fail on the box but a program that fails it never needed a paid attempt. Read the catalog's band for the endpoint you are about to call, compare the estimate with GET /v1/balance, and stop before a 402. After the run, read the usage summary and confirm final is true before you treat the job as settled.
What is not covered: there is no top-up endpoint. The docs say top-ups are a dashboard operation, with the public API exposing balance and usage reads only.
A simple pattern follows from this: poll status as often as your rate budget allows, read the balance before each large batch, and treat a 402 as a stop signal rather than something to retry. Retrying a 402 without adding funds will return the same answer, while a 429 is worth retrying after the retry-after delay.
One more reason to keep the free calls in your loop: they give you the numbers a finance or ops reviewer will ask for. The balance read shows what is left, the usage read shows what each job debited, and the catalog shows the band you should expect before you submit. None of them changes the wallet, so you can run them as often as the rate budget allows without adding to the bill.
Sources
Related posts
More in Pricing
- How Sume pricing works: plans, one wallet, published model rates
Sume plans set access and concurrency. Usage draws from one prepaid wallet at each model's published USD rate, for generation, the Agent, Formats, and the API.
- AI avatar video API pricing: cost per second and per minute
Sume bills AI avatar video per second by quality tier, with separate rates when you send a product image. Per-minute costs for standard, plus, and max.
- AI video generation cost per video: what one Sume API run cost
To see what one Sume run or video job cost, call GET /v1/usage with run_id or job_id and read debited_usd: the wallet deduction, agent turns included.
- Estimate AI video generation cost before running a Sume job
See what an AI video will cost on Sume before paying: published rates, GET /v1/catalog estimates, unbilled plan checks, MCP dry runs, and spend caps.
Written by Sume