Failed AI video job on Sume: is the reservation released?

Sume reserves the estimated cost at submit and captures it on success; failed jobs release or refund it where applicable. How to check your own account.

5 min readSume
All posts

Sume reserves the estimated cost when it accepts a video job, captures it when the job completes, and where applicable releases or refunds the reservation for a failed job or a failed queue admission. That wording is from the Sume generation admission docs (read 2026-10-06). "Where applicable" is the docs' own hedge, so do not assume every failure refunds in full; the outcome shows in your own usage.

This question is common after the OpenAI Sora API ended on 2026-09-24 (Magic Hour tracker, read 2026-10-06), because teams comparing vendors want to know what a bad render costs. The honest answer is to measure it on your own jobs.

What the docs give you

Billing states for a video job, Sume docs, read 2026-10-06
MomentWhat Sume doesHow you see it
Submit acceptedReserves the billable estimate (list times 1.25)The job exists; balance is reduced by the hold
Submit refused with 402No job, no provider work402 insufficient_credits
Job completesCaptures the reserved usageusage.cost on the poll response
Job failsReleases or refunds the reservation where applicablestatus failed; check usage
Cancelled before it startsCancellation succeeds only before generation startsstatus cancelled or canceled, by route

Measure it yourself

Do not take a table's word for your own account. Log usage.cost for every terminal job, including failed ones, and total the failed ones. A small script over your stored poll responses does it.

import json

totals = {"completed": 0.0, "failed": 0.0, "canceled": 0.0}
ALIAS = {"cancelled": "canceled"}
counts = dict.fromkeys(totals, 0)
for line in open("polls.jsonl"):
    j = json.loads(line)
    s = ALIAS.get(j.get("status"), j.get("status"))
    if s not in totals:
        continue
    counts[s] += 1
    totals[s] += float((j.get("usage") or {}).get("cost") or 0)
for s in totals:
    print(s, counts[s], round(totals[s], 2))
if counts["failed"] and totals["failed"] > 0:
    print("failed jobs carried cost: check usage with support")

What to do with a failed job

  • Read the error field before you retry. A prompt the model refuses will fail again, and a transient provider error may not.
  • Retry with the same Idempotency-Key only when the request is unchanged; a changed prompt needs a new key.
  • Do not resubmit a job that is merely slow. A running job still bills when it completes.
  • Cancel queued jobs you no longer need; cancel works only before generation starts, and after that the API answers 409 job_generation_already_started.

A balance that is too low is a different case: the 402 arrives at submit, before any job exists. The post on checking the balance before a back-catalog run covers that, and the usage.cost logging post shows how to keep the ledger the script above reads.

If a failed job shows a cost you cannot explain, keep the job id and the events from GET /v1/jobs/{id}/events and send both with your question; they are the public timeline of what happened.

A small ledger that answers it for your account

Three columns are enough: job id, terminal status, and usage.cost. After a week of real jobs, filter to failed and look at the third column. If it is empty or zero for failures, the reservation was released in your case. If it is not, you have a concrete job id to bring to support, with the events timeline attached.

Keep the check separate from retry logic. The ledger answers a finance question, not an engineering one, and mixing them leads to retry policies that are tuned to save a cent rather than to ship the clip. Retry on the error category, as a policy, and audit the cost afterwards.

If your finance team asks for a rule of thumb, give them the measured figure from your own ledger, with the dates it covers, and not a sentence from a blog. The first month's numbers will be small, but they are yours, and they get better each month as the sample grows.

  • Store the raw poll response for every terminal job; it has the status, error and usage you need.
  • Record the time of submit and the time of the terminal poll, so a refund question can say how long the job ran.
  • Do not retry a failed job with the same key and an unchanged request in a tight loop; the replay returns the original failed job.
  • Re-read the admission docs when you plan a quarter's budget, since the wording about release and refund is the vendor's to change.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume