Balance needed for one AI generation call: published max per endpoint
Sume's catalog publishes estimated, minimum and maximum cents per endpoint. The worst-case hold for one call runs from 3 cents to $56.25 before a 402 can fire.

The balance you need for one Sume generation call is the estimate for that request, and the catalog publishes a band around it: estimated_usd_cents, minimum_usd_cents and maximum_usd_cents for each endpoint under model_pricing in GET /v1/catalog. The worst-case column runs from 3 cents for a background removal to $56.25 for a 300-second VEED Fabric clip at 720p. If the workspace cannot cover the estimate for your exact request, the call returns 402 insufficient_credits and nothing starts.
This matters most for batches, because a wallet that covers one cheap call may not cover the next expensive one. Below are the published bands, what a band means, and a short script that checks the wallet against a request you have already priced.
What does the catalog band mean?
Sume's catalog describes its reservation policy in one sentence: submit-time reservation stores the estimated billable amount in USD micros, rounds up to cents for balance compatibility, captures on completion, and refunds failed or pre-generation canceled jobs before capture. The estimate follows the fields you send, such as the duration, resolution or character count.
The band is the range the estimate can take across valid requests for that endpoint. Minimum is the cheapest valid request, estimated is what an absent-field default request reserves, and maximum is the largest request the endpoint accepts. Your request lands somewhere inside it, and you can compute the point from the per-unit rate on the pricing page.
What is the worst case for each endpoint?
The figures below are copied from Sume's generated catalog constants, in US cents, so they match what the catalog publishes.
| Endpoint | Estimated | Minimum | Maximum | What sets the maximum |
|---|---|---|---|---|
| Remove background | 3 | 3 | 3 | Flat per image |
| Speech transcription | 1 | 1 | 10 | 10 minutes of audio |
| Text to speech | 5 | 1 | 95 | 20,000 characters |
| Music generation | 13 | 13 | 13 | Flat per track |
| Image upscale | 20 | 20 | 20 | Flat per image |
| Video upscale | 5 | 1 | 27 | 30 seconds of input |
| Timeline render | 10 | 10 | 300 | 30 output minutes |
| MiniMax H3 Max lip-sync | 50 | 32 | 300 | 14.8 s (15 billed) at 1080p |
| Kling 3.0 Motion Control | 79 | 16 | 473 | 30 s of motion video |
| Image 1.0 | 23 | 1 | 736 | Max quality, 4 images, 16 references |
| Video 1.0 | 232 | 57 | 1742 | Longest 1080p or 720p lane at 21:9 |
| VEED Fabric 1.0 | 94 | 10 | 5625 | 300 s at 720p |
Why does the maximum matter if I send a small request?
It usually does not, which is the point of reading the band. A 10-minute STT upload can never hold more than 10 cents, and a flat-priced endpoint holds the same amount every time. The risk is endpoints where one field drives the price. For Fabric, a 5 second request reserves around 94 cents, but the same call with duration_seconds: 300 reserves $56.25. A bulk script that copies one request body across many jobs can therefore exhaust a balance far faster than its first job suggests.
The fix is to compute each reservation from the unit rate and your own duration before you submit, then compare the sum of in-flight jobs to the balance. A 429 queue_full also releases the hold for the failed attempt, per the generation admission docs, so a rejected submit does not leave money tied up.
- Flat endpoints (background removal, music, image upscale): the band is one number.
- Per-second endpoints: reservation is seconds times rate, so the field you send is the cost.
- Image 1.0 and Video 1.0: the maximum assumes the largest request shape, so a normal request is far below it.
- Reads such as
GET /v1/balance,GET /v1/usageand the catalog do not reserve anything.
How do I check the wallet before a batch?
GET /v1/balance returns data.balance.available_amount_usd_micros, the spendable USD in millionths of a dollar. Compare it with the sum of the holds you are about to create. This script prices a list of Fabric clips from their audio seconds at the published 720p rate of $0.1875 per second and stops if the wallet is short:
If you run several jobs at once, add the holds together, because each accepted job reserves its own estimate and the balance is checked at each submit. A wallet that covers one job can still return a 402 on the third.
In practice, keep a floor in your own code that is a little above the largest figure on the list you actually call, and refill before you cross it. That turns a surprise 402 mid-batch into a planned stop, and it costs nothing, because the balance read does not move money. Treat a 402 as a refusal to start rather than a failed job, since no hold is placed.
import math, os, requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
RATE_MICROS = 187_500 # VEED Fabric 1.0 720p, per audio second
def available_micros():
r = requests.get("https://api.sume.com/v1/balance", headers=H, timeout=30)
r.raise_for_status()
return r.json()["data"]["balance"]["available_amount_usd_micros"]
def need_micros(seconds_list):
return sum(math.ceil(s) * RATE_MICROS for s in seconds_list)
clips = [12.4, 30.0, 61.5]
need, have = need_micros(clips), available_micros()
print(f"need ${need/1e6:.4f}, have ${have/1e6:.4f}")
if need > have:
raise SystemExit("top up or shorten the clips before submitting")Sources
Related posts
More in Pricing
- Real-time AI avatar pricing: live minutes vs one render
Live avatars bill by conversation minute and concurrent stream; a rendered clip is paid once and watched by anyone. Tavus plan numbers and the break-even.
- Sume plans: Pro $40, Startup $120, Scale $400 and what they limit
What each Sume plan sets: monthly price, concurrent jobs, queue size and API write budget. Usage is billed at each model's rate, not by the plan.
- TTS cost by character: the same sentence in four languages
Sume bills TTS per character, not per second. One sentence counted in English, French, German and Korean, with the 1-cent floor and the 20,000-character cap.
- Which Sume API calls are free: balance, usage, catalog, filter check
Balance, usage and catalog reads cost nothing on Sume, and neither does the video-filter check. What is billed, what is refunded, and what a 402 means.
Written by Sume