Check balance before bulk transcription: 402 and admission preview

A 402 insufficient_credits arrives before any provider work. Read GET /v1/balance and POST /v1/generation/admission-preview first and size the run.

5 min readSume
All posts

Before a bulk transcription run on Sume, read GET /v1/balance for the wallet and POST /v1/generation/admission-preview for queue room and request cost. Both are read-only: the preview does not create a job, reserve credits or call a provider. If you skip this and the balance cannot cover an estimate, the submit answers 402 insufficient_credits before provider work starts. Facts from Generation admission and the OpenAPI route descriptions, read 2026-10-06.

What does the 402 mean for a batch?

Sume reserves the estimated amount at submit time. For STT 1.0 the estimate is prorated by second from duration_seconds, and one minute ($0.01) when the field is omitted. A batch of 500 clips with no hint therefore needs $5.00 of headroom for the jobs in flight, even if the clips are 6 seconds long.

The docs tell you what to do on a 402: upgrade the plan, wait for included Gen$, or submit a cheaper request. They also say not to invent prepaid top-ups.

Admission controls that can stop a submit, from Generation admission (docs.sume.com), read 2026-10-06.
Status and codeCauseClient behavior
402 insufficient_creditsThe estimate cannot be reserved from the balanceShrink the request or add funds; do not retry unchanged
429 queue_fullConcurrency and queue are both fullWait or cancel queued jobs; retry with the same key
429 rate_limitedRequest volume over the abuse limitBack off with retry-after

A preflight in two calls

The preview body takes model (the public id, sume/stt-1.0) and request, the body you intend to submit. Omit request to inspect balance and queue only.

import os, requests

API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

balance = requests.get(f"{API}/v1/balance", headers=H, timeout=30)
balance.raise_for_status()
print(balance.json()["data"]["balance"])

preview = requests.post(
    f"{API}/v1/generation/admission-preview",
    headers=H,
    json={
        "model": "sume/stt-1.0",
        "request": {"audio_url": "https://media.sume.com/example.wav", "duration_seconds": 8},
    },
    timeout=30,
)
preview.raise_for_status()
print(preview.json())

How big a wave can I submit?

Use the headroom rule from the docs: max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs), capped at queue_capacity_remaining. At zero headroom, wait and refresh instead of submitting. The per-plan numbers are in transcribing 10,000 short clips.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume