Check balance before bulk transcription: 402 and admission preview
A 402 insufficient_credits arrives before any provider work. Read GET /v1/balance and POST /v1/generation/admission-preview first and size the run.

Before a bulk transcription run on Sume, read GET /v1/balance for the wallet and POST /v1/generation/admission-preview for queue room and request cost. Both are read-only: the preview does not create a job, reserve credits or call a provider. If you skip this and the balance cannot cover an estimate, the submit answers 402 insufficient_credits before provider work starts. Facts from Generation admission and the OpenAPI route descriptions, read 2026-10-06.
What does the 402 mean for a batch?
Sume reserves the estimated amount at submit time. For STT 1.0 the estimate is prorated by second from duration_seconds, and one minute ($0.01) when the field is omitted. A batch of 500 clips with no hint therefore needs $5.00 of headroom for the jobs in flight, even if the clips are 6 seconds long.
The docs tell you what to do on a 402: upgrade the plan, wait for included Gen$, or submit a cheaper request. They also say not to invent prepaid top-ups.
| Status and code | Cause | Client behavior |
|---|---|---|
402 insufficient_credits | The estimate cannot be reserved from the balance | Shrink the request or add funds; do not retry unchanged |
429 queue_full | Concurrency and queue are both full | Wait or cancel queued jobs; retry with the same key |
429 rate_limited | Request volume over the abuse limit | Back off with retry-after |
A preflight in two calls
The preview body takes model (the public id, sume/stt-1.0) and request, the body you intend to submit. Omit request to inspect balance and queue only.
import os, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
balance = requests.get(f"{API}/v1/balance", headers=H, timeout=30)
balance.raise_for_status()
print(balance.json()["data"]["balance"])
preview = requests.post(
f"{API}/v1/generation/admission-preview",
headers=H,
json={
"model": "sume/stt-1.0",
"request": {"audio_url": "https://media.sume.com/example.wav", "duration_seconds": 8},
},
timeout=30,
)
preview.raise_for_status()
print(preview.json())How big a wave can I submit?
Use the headroom rule from the docs: max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs), capped at queue_capacity_remaining. At zero headroom, wait and refresh instead of submitting. The per-plan numbers are in transcribing 10,000 short clips.
Sources
Related posts
More in Developers
- Port a Bedrock image call to Sume /v1/images in Python
Moving from boto3 invoke_model for Nova Canvas or Titan to Sume's REST call: the request mapping, the response shape, the status codes, and the swap code.
- BullMQ delayed job that polls an AI video job and reschedules itself
A BullMQ worker reads Sume's job status once, then adds the next poll with a delay from next_poll_after_seconds, so no worker slot is held while a clip renders.
- A calendar file for AI model shutdown dates: .ics from Python
Generate an .ics file with all-day events and 14-day reminders for the gpt-image-1 and gpt-image-1.5 shutdown dates, then import it into any calendar app.
- Cancel a queued transcription job: what the 409 means
POST /v1/jobs/{id}/cancel works only before work starts. After that it returns 409 job_generation_already_started and the job finishes and bills normally.
Written by Sume