Balance needed for one AI generation call: published max per endpoint

Sume's catalog publishes estimated, minimum and maximum cents per endpoint. The worst-case hold for one call runs from 3 cents to $56.25 before a 402 can fire.

5 min readSume
All posts

The balance you need for one Sume generation call is the estimate for that request, and the catalog publishes a band around it: estimated_usd_cents, minimum_usd_cents and maximum_usd_cents for each endpoint under model_pricing in GET /v1/catalog. The worst-case column runs from 3 cents for a background removal to $56.25 for a 300-second VEED Fabric clip at 720p. If the workspace cannot cover the estimate for your exact request, the call returns 402 insufficient_credits and nothing starts.

This matters most for batches, because a wallet that covers one cheap call may not cover the next expensive one. Below are the published bands, what a band means, and a short script that checks the wallet against a request you have already priced.

What does the catalog band mean?

Sume's catalog describes its reservation policy in one sentence: submit-time reservation stores the estimated billable amount in USD micros, rounds up to cents for balance compatibility, captures on completion, and refunds failed or pre-generation canceled jobs before capture. The estimate follows the fields you send, such as the duration, resolution or character count.

The band is the range the estimate can take across valid requests for that endpoint. Minimum is the cheapest valid request, estimated is what an absent-field default request reserves, and maximum is the largest request the endpoint accepts. Your request lands somewhere inside it, and you can compute the point from the per-unit rate on the pricing page.

What is the worst case for each endpoint?

The figures below are copied from Sume's generated catalog constants, in US cents, so they match what the catalog publishes.

Published catalog bands in USD cents, from Sume's catalog constants, read 2026-10-03
EndpointEstimatedMinimumMaximumWhat sets the maximum
Remove background333Flat per image
Speech transcription111010 minutes of audio
Text to speech519520,000 characters
Music generation131313Flat per track
Image upscale202020Flat per image
Video upscale512730 seconds of input
Timeline render101030030 output minutes
MiniMax H3 Max lip-sync503230014.8 s (15 billed) at 1080p
Kling 3.0 Motion Control791647330 s of motion video
Image 1.0231736Max quality, 4 images, 16 references
Video 1.0232571742Longest 1080p or 720p lane at 21:9
VEED Fabric 1.094105625300 s at 720p

Why does the maximum matter if I send a small request?

It usually does not, which is the point of reading the band. A 10-minute STT upload can never hold more than 10 cents, and a flat-priced endpoint holds the same amount every time. The risk is endpoints where one field drives the price. For Fabric, a 5 second request reserves around 94 cents, but the same call with duration_seconds: 300 reserves $56.25. A bulk script that copies one request body across many jobs can therefore exhaust a balance far faster than its first job suggests.

The fix is to compute each reservation from the unit rate and your own duration before you submit, then compare the sum of in-flight jobs to the balance. A 429 queue_full also releases the hold for the failed attempt, per the generation admission docs, so a rejected submit does not leave money tied up.

  • Flat endpoints (background removal, music, image upscale): the band is one number.
  • Per-second endpoints: reservation is seconds times rate, so the field you send is the cost.
  • Image 1.0 and Video 1.0: the maximum assumes the largest request shape, so a normal request is far below it.
  • Reads such as GET /v1/balance, GET /v1/usage and the catalog do not reserve anything.

How do I check the wallet before a batch?

GET /v1/balance returns data.balance.available_amount_usd_micros, the spendable USD in millionths of a dollar. Compare it with the sum of the holds you are about to create. This script prices a list of Fabric clips from their audio seconds at the published 720p rate of $0.1875 per second and stops if the wallet is short:

If you run several jobs at once, add the holds together, because each accepted job reserves its own estimate and the balance is checked at each submit. A wallet that covers one job can still return a 402 on the third.

In practice, keep a floor in your own code that is a little above the largest figure on the list you actually call, and refill before you cross it. That turns a surprise 402 mid-batch into a planned stop, and it costs nothing, because the balance read does not move money. Treat a 402 as a refusal to start rather than a failed job, since no hold is placed.

import math, os, requests

H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
RATE_MICROS = 187_500  # VEED Fabric 1.0 720p, per audio second

def available_micros():
    r = requests.get("https://api.sume.com/v1/balance", headers=H, timeout=30)
    r.raise_for_status()
    return r.json()["data"]["balance"]["available_amount_usd_micros"]

def need_micros(seconds_list):
    return sum(math.ceil(s) * RATE_MICROS for s in seconds_list)

clips = [12.4, 30.0, 61.5]
need, have = need_micros(clips), available_micros()
print(f"need ${need/1e6:.4f}, have ${have/1e6:.4f}")
if need > have:
    raise SystemExit("top up or shorten the clips before submitting")

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume