Kill switch for a Sume video batch: what it stops and what not
A stop file checked between Sume submits halts a batch, but jobs already accepted keep running and billing. A tested Python loop and the cancel limits.

A simple cost guard is a file you check before every submit: touch it and the batch stops creating jobs. It does nothing for jobs already accepted, which keep running and billing, and cancel only works before the provider has started the generation.
The loop below shows the pattern with a placeholder submit function. In your code, submit should send an Idempotency-Key and return the job id.
The loop
It stores each job id before moving on and stops when the file appears. The test run stopped at the third shot and kept the first three ids.
Placing the check just before the submit, and the write of the job id just after it, keeps the window small. If the process dies between the two, you may have a job you have no id for, which is exactly where the Idempotency-Key helps: on restart, submit that shot again with the same key and the same payload, and you get the original job back instead of a second paid one. Keep the key stable by deriving it from your shot id, not from a timestamp or a random value generated at call time.
import pathlib, time
STOP = pathlib.Path("STOP_SUBMITS") # touch this file to halt the batch
def run_batch(shots, submit, store):
"""submit(shot) -> job id (send an Idempotency-Key inside it); store(shot, job_id) persists."""
for shot in shots:
if STOP.exists():
print("stop file found; not submitting", shot["id"])
break
job_id = submit(shot) # a 2xx means the job exists and is reserved
store(shot, job_id) # persist before the next iteration
time.sleep(0.2)
# Jobs already submitted keep running and keep billing. Cancel the
# queued ones explicitly (POST /v1/jobs/{id}/cancel) if you want them gone.
if __name__ == "__main__":
shots = [{"id": i} for i in range(5)]
done = []
def fake_submit(s):
if s["id"] == 2: STOP.touch()
return f"job_{s['id']}"
run_batch(shots, fake_submit, lambda s, j: done.append(j))
print(done); STOP.unlink()What the guard cannot do
Per Sume's docs, credits are reserved at submit, captured on completion and refunded on failure. A job you have already submitted has its reservation. Cancel is available only before provider submission starts; after that it returns 409 job_generation_already_started.
| Job state | Effect of stopping |
|---|---|
| Not submitted yet | Never created, no charge |
| Queued, provider not started | Still runs unless you cancel with POST /v1/jobs/{id}/cancel |
| Provider generation started | Cannot be canceled; completes and bills |
| Completed | Already captured |
Pair it with plan limits
Generation admission also caps what you can have in flight. Free allows 1 processing and 5 queued, Pro 4 and 20, Startup 8 and 40, Scale 20 and 100, so a runaway loop hits queue_full (429) before it hits unlimited spend.
Tradeoffs
A file check is crude but works across processes and needs no service. It does not roll back anything, so keep the stored ids and cancel the queued ones deliberately if you want them gone.
Before you run any large batch, decide who may create the stop file and where it lives, and write that in the runbook. A guard nobody knows how to trigger at three in the morning is not much of a guard, and a shared network path is more reliable than a file on a laptop that may be asleep.
Sources
Related posts
More in Developers
- Ktor webhook route for an AI video job: receiveText and HMAC SHA-256
A Ktor route reads the raw body with call.receiveText(), checks Sume's sume-v1 HMAC over timestamp.body in Kotlin and returns 401 when the secret is empty.
- Label speakers without diarization: transcribe each mic track, merge
Sume STT has no speaker labels. If you record each speaker on a separate track, transcribe both tracks and merge the segments by start time. Python script.
- Laravel queued job for an AI video API: delay, redispatch, poll
A Laravel ShouldQueue job reads a Sume video job once and redispatches itself with ->delay() from next_poll_after_seconds, so no worker sleeps during a render.
- Let QA override the image model per request, with an allowlist
An internal render endpoint that honors an X-Image-Model header only for allowlisted Sume ids, so QA can test gpt-image-2.5 before the config flips. Python.
Written by Sume