Kill switch for a Sume video batch: what it stops and what not

A stop file checked between Sume submits halts a batch, but jobs already accepted keep running and billing. A tested Python loop and the cancel limits.

4 min readSume
All posts

A simple cost guard is a file you check before every submit: touch it and the batch stops creating jobs. It does nothing for jobs already accepted, which keep running and billing, and cancel only works before the provider has started the generation.

The loop below shows the pattern with a placeholder submit function. In your code, submit should send an Idempotency-Key and return the job id.

The loop

It stores each job id before moving on and stops when the file appears. The test run stopped at the third shot and kept the first three ids.

Placing the check just before the submit, and the write of the job id just after it, keeps the window small. If the process dies between the two, you may have a job you have no id for, which is exactly where the Idempotency-Key helps: on restart, submit that shot again with the same key and the same payload, and you get the original job back instead of a second paid one. Keep the key stable by deriving it from your shot id, not from a timestamp or a random value generated at call time.

import pathlib, time

STOP = pathlib.Path("STOP_SUBMITS")      # touch this file to halt the batch

def run_batch(shots, submit, store):
    """submit(shot) -> job id (send an Idempotency-Key inside it); store(shot, job_id) persists."""
    for shot in shots:
        if STOP.exists():
            print("stop file found; not submitting", shot["id"])
            break
        job_id = submit(shot)            # a 2xx means the job exists and is reserved
        store(shot, job_id)              # persist before the next iteration
        time.sleep(0.2)
    # Jobs already submitted keep running and keep billing. Cancel the
    # queued ones explicitly (POST /v1/jobs/{id}/cancel) if you want them gone.

if __name__ == "__main__":
    shots = [{"id": i} for i in range(5)]
    done = []
    def fake_submit(s):
        if s["id"] == 2: STOP.touch()
        return f"job_{s['id']}"
    run_batch(shots, fake_submit, lambda s, j: done.append(j))
    print(done); STOP.unlink()

What the guard cannot do

Per Sume's docs, credits are reserved at submit, captured on completion and refunded on failure. A job you have already submitted has its reservation. Cancel is available only before provider submission starts; after that it returns 409 job_generation_already_started.

What stops when you stop the loop (read 2026-10-06, Sume docs)
Job stateEffect of stopping
Not submitted yetNever created, no charge
Queued, provider not startedStill runs unless you cancel with POST /v1/jobs/{id}/cancel
Provider generation startedCannot be canceled; completes and bills
CompletedAlready captured

Pair it with plan limits

Generation admission also caps what you can have in flight. Free allows 1 processing and 5 queued, Pro 4 and 20, Startup 8 and 40, Scale 20 and 100, so a runaway loop hits queue_full (429) before it hits unlimited spend.

Tradeoffs

A file check is crude but works across processes and needs no service. It does not roll back anything, so keep the stored ids and cancel the queued ones deliberately if you want them gone.

Before you run any large batch, decide who may create the stop file and where it lives, and write that in the runbook. A guard nobody knows how to trigger at three in the morning is not much of a guard, and a shared network path is more reliable than a file on a laptop that may be asleep.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume