Luma API callbacks and credit balance vs Sume webhooks and /v1/balance
Luma's docs list callbacks and a credits balance. Sume has webhook mode, GET /v1/balance, and a generation_limits snapshot. How to use each before a batch.

Both a callback and a way to read your remaining credit are table stakes for a generation API, and Sume has both: mode: "webhook" for push, GET /v1/balance for a USD balance, and a generation_limits snapshot in every submit response. The Luma API changelog lists topics for callbacks and a credits balance, so a client that wraps both services can use the same shape for each.
The point of the balance read is to decide before a batch, not after the first failure.
What the changelog lists
The changelog topics, as listed.
| Topic | What the page lists |
|---|---|
| Upscale | Up to 4K, with progress |
| Models | Photon and Photon Flash |
| Callbacks | Listed |
| Credits | A balance read |
The Sume equivalents
Sume's balance is USD-denominated. GET /v1/balance returns it, and compatibility fields can expose rounded cent values as credits. GET /v1/usage returns the ledger, where each row is reserved, captured or refunded: a reservation is held when a job is accepted, captured on success, and released after failure or cancellation before capture.
| Question | Read | Notes |
|---|---|---|
| How much can I spend? | GET /v1/balance | USD-denominated |
| What did one job cost? | GET /v1/usage?job_id=... | Also run_id and thread_id |
| How much can I queue? | generation_limits in a submit response | A snapshot; can change right after |
| Did the push arrive? | Job object and job events | Delivery status and attempt count |
Before a batch
Read the balance, then read generation_limits from your first submit. Use queue_capacity_remaining to decide how many more jobs to add, and treat wave_size_hint as a hint only, since it is not a concurrency limit. A workspace that cannot reserve an estimate gets 402 insufficient_credits before provider work starts, and one with no accepted capacity gets 429 queue_full.
Neither error is a reason to loop faster. Wait, or cancel queued jobs you no longer need, then retry with the same idempotency key.
Callbacks on Sume
The push contract is narrow and worth stating exactly.
- Send
mode: "webhook"with a public HTTPSwebhook_url, or sendwebhook_urlalone and the mode is inferred. - Events are terminal only:
job.completed,job.failed,job.canceled. There are no progress callbacks. - Verify
x-sume-webhook-signatureagainst the raw body, and usejob_idas your idempotency key. - If you need progress, read
GET /v1/jobs/{id}/events, a pull snapshot, not a stream.
A small preflight
The script reads the balance and refuses to start if it is below an amount you choose. It does not create any job. Set the threshold from your own estimate, such as the sum of dry_run previews, and note that the field names should be checked against the live OpenAPI schema before you depend on them.
const res = await fetch("https://api.sume.com/v1/balance", {
headers: { Authorization: `Bearer ${process.env.SUME_API_KEY}` },
});
if (!res.ok) throw new Error(`balance read failed: ${res.status}`);
const balance = await res.json();
console.log(JSON.stringify(balance));
// Compare the USD field in the live schema with your batch estimate here
// and stop before the first paid submit if it does not cover it.Where to look for the exact fields
Sume's docs point to the live OpenAPI schema at https://api.sume.com/reference/json for exact response fields. Use it rather than copying names from a blog post, including this one.
Sources
Related posts
More in Comparisons
- Luma Ray 2 API parameters: keyframes, loop and callback_url
Luma's video docs list ray-2 and ray-flash-2 with keyframes, loop, concepts and callback_url. Here is each field mapped to the Sume /v1/videos request.
- Midjourney has no public API: comparable control in a pipeline
Midjourney's 10/1 alpha adds a pinnable --exp setting and edits that keep aspect ratio. What to use when you need that kind of control from code.
- MiniMax H3 vs H3-Max limits: the MiniMax page next to Sume's catalog
MiniMax lists H3 at 4 to 15 s and H3-Max at 5 to 15 s with different resolutions. Sume's catalog states its own ranges. Side by side, with the file-size limits.
- Pika API Club: 100+ models at reduced pricing, questions to ask first
Pika's API Club (Aug 5) offers 100+ models at reduced pricing, and the Sep 17 relaunch adds audio models. Seven questions to ask an aggregator.
Written by Sume