Failed AI video job on Sume: is the reservation released?
Sume reserves the estimated cost at submit and captures it on success; failed jobs release or refund it where applicable. How to check your own account.

Sume reserves the estimated cost when it accepts a video job, captures it when the job completes, and where applicable releases or refunds the reservation for a failed job or a failed queue admission. That wording is from the Sume generation admission docs (read 2026-10-06). "Where applicable" is the docs' own hedge, so do not assume every failure refunds in full; the outcome shows in your own usage.
This question is common after the OpenAI Sora API ended on 2026-09-24 (Magic Hour tracker, read 2026-10-06), because teams comparing vendors want to know what a bad render costs. The honest answer is to measure it on your own jobs.
What the docs give you
| Moment | What Sume does | How you see it |
|---|---|---|
| Submit accepted | Reserves the billable estimate (list times 1.25) | The job exists; balance is reduced by the hold |
| Submit refused with 402 | No job, no provider work | 402 insufficient_credits |
| Job completes | Captures the reserved usage | usage.cost on the poll response |
| Job fails | Releases or refunds the reservation where applicable | status failed; check usage |
| Cancelled before it starts | Cancellation succeeds only before generation starts | status cancelled or canceled, by route |
Measure it yourself
Do not take a table's word for your own account. Log usage.cost for every terminal job, including failed ones, and total the failed ones. A small script over your stored poll responses does it.
import json
totals = {"completed": 0.0, "failed": 0.0, "canceled": 0.0}
ALIAS = {"cancelled": "canceled"}
counts = dict.fromkeys(totals, 0)
for line in open("polls.jsonl"):
j = json.loads(line)
s = ALIAS.get(j.get("status"), j.get("status"))
if s not in totals:
continue
counts[s] += 1
totals[s] += float((j.get("usage") or {}).get("cost") or 0)
for s in totals:
print(s, counts[s], round(totals[s], 2))
if counts["failed"] and totals["failed"] > 0:
print("failed jobs carried cost: check usage with support")What to do with a failed job
- Read the
errorfield before you retry. A prompt the model refuses will fail again, and a transient provider error may not. - Retry with the same
Idempotency-Keyonly when the request is unchanged; a changed prompt needs a new key. - Do not resubmit a job that is merely slow. A running job still bills when it completes.
- Cancel queued jobs you no longer need; cancel works only before generation starts, and after that the API answers
409 job_generation_already_started.
A balance that is too low is a different case: the 402 arrives at submit, before any job exists. The post on checking the balance before a back-catalog run covers that, and the usage.cost logging post shows how to keep the ledger the script above reads.
If a failed job shows a cost you cannot explain, keep the job id and the events from GET /v1/jobs/{id}/events and send both with your question; they are the public timeline of what happened.
A small ledger that answers it for your account
Three columns are enough: job id, terminal status, and usage.cost. After a week of real jobs, filter to failed and look at the third column. If it is empty or zero for failures, the reservation was released in your case. If it is not, you have a concrete job id to bring to support, with the events timeline attached.
Keep the check separate from retry logic. The ledger answers a finance question, not an engineering one, and mixing them leads to retry policies that are tuned to save a cent rather than to ship the clip. Retry on the error category, as a policy, and audit the cost afterwards.
If your finance team asks for a rule of thumb, give them the measured figure from your own ledger, with the dates it covers, and not a sentence from a blog. The first month's numbers will be small, but they are yours, and they get better each month as the sample grows.
- Store the raw poll response for every terminal job; it has the
status,errorandusageyou need. - Record the time of submit and the time of the terminal poll, so a refund question can say how long the job ran.
- Do not retry a failed job with the same key and an unchanged request in a tight loop; the replay returns the original failed job.
- Re-read the admission docs when you plan a quarter's budget, since the wording about release and refund is the vendor's to change.
Sources
Related posts
More in Pricing
- Cost per second of AI video on Sume: Omni, Wan and H3 Max priced
A per-second and per-clip price table for the video models Sume lists with a fal list price, after the Sora API ended, with the 1.25 multiple already applied.
- Sume TTS vs MAI-Voice-2.1 at 50k to 5M characters a month
At 1M characters a month Sume TTS bills $47.50, MAI-Voice-2.1 lists $22 and Flash $15. The monthly table for 50k to 5M and when the gap is worth paying.
- TTS API price per 1,000 characters: Sume, MAI, Groq, Lemonfox, Unreal
A per-1,000-character table of list prices for Sume, MAI-Voice-2.1 and Flash, Groq Orpheus, Lemonfox and Unreal Speech, with cost for a 750-character short.
- Transcribe 100,000 two-second voice command recordings: the bill
A two-second clip costs about $0.000333 on Sume STT, so 100,000 cost $33.33. Without duration_seconds each job in flight holds a full $0.01 instead.
Written by Sume