systemd restart: finish Sume jobs with sume jobs watch, never resubmit
A systemd unit that restarts a finisher and resumes stored Sume job ids with sume jobs watch and download, so a crash never pays for a second job.

Write each job_id to disk the moment a submit returns, and let a systemd service that restarts on failure walk that directory. For each id it runs sume jobs watch <job_id>, then sume jobs download <job_id> --output-dir ./out, and removes the id file only after the download succeeds. A crash, a reboot or a systemctl restart then resumes the same jobs, and nothing submits a second paid request.
This is the recovery rule from the CLI jobs docs: if a process restarts or a local timeout occurs, use the job recovery commands and do not submit paid generation again. jobs watch polls until the job is terminal or a timeout occurs, and jobs download writes completed media artifacts to a directory. The helpers also work for Image, Video and Music jobs created through the Developer API.
The unit
The unit runs as a long-lived service and restarts after a failure. The API key comes from an environment file, using the SUME_API_KEY variable that the CLI reads.
[Unit]
Description=Finish pending Sume jobs
After=network-online.target
Wants=network-online.target
[Service]
User=sume
EnvironmentFile=/etc/sume/env
ExecStart=/usr/local/bin/finish-sume-jobs.sh
Restart=on-failure
RestartSec=30
[Install]
WantedBy=multi-user.targetThe script
The script exits non-zero if any job fails to finish cleanly, which makes systemd restart it after RestartSec. A job that ended failed or canceled is a final answer: record it and remove the id so it is not retried.
#!/usr/bin/env bash
set -u
: "${SUME_API_KEY:?SUME_API_KEY is empty}"
PENDING=/var/lib/sume/pending
OUT=/var/lib/sume/out
mkdir -p "$PENDING" "$OUT"
rc=0
for f in "$PENDING"/*.id; do
[ -e "$f" ] || continue
id=$(cat "$f")
if sume jobs watch "$id" && sume jobs download "$id" --output-dir "$OUT/$id"; then
rm -f "$f"
else
echo "job $id not finished or not downloadable yet" >&2
rc=1
fi
done
exit $rcWhich command handles which situation
Match each situation to the command that handles it. None of them is a new submit.
| Situation | Command |
|---|---|
| Process restarted, job still running | sume jobs watch <job_id> |
| Need the current state once | sume jobs status <job_id> --agent --json |
| Job completed, files needed locally | sume jobs download <job_id> --output-dir ./out |
| Want the timeline of a slow job | sume jobs events <job_id> --agent --json |
| Queued job you no longer want | sume jobs cancel <job_id> --confirm-submit |
Caveats
- The CLI does not ship
sume videoorsume imagegenerators. Submit those jobs through the HTTP API and use the CLI only to recover them. - Cancel works only before generation starts. After that the API answers
409 job_generation_already_startedand the job finishes normally. - A local timeout does not cancel anything. The job keeps running and keeps billing, so keep the id file until the download succeeds.
Sources
Related posts
More in Developers
- End-user id on jobs: OpenAI safety identifier vs Sume metadata
OpenAI's Realtime guide asks for an OpenAI-Safety-Identifier header. Sume stores caller metadata on the job but does not send it to the provider. Use both.
- Temporal Paygo starter: submit and poll a Sume job
Temporal's Paygo plan has a $0 monthly minimum. A first workflow can submit a Sume job in one Activity and poll its status in a second. Python code included.
- Test a faster-and-cheaper claim with your own timings and usage.cost
Luma's news page says Ray3.14 is 4x faster and 3x cheaper. A short Python script turns your own job timings and usage.cost values into two ratios you can trust.
- Test MAI-Transcribe-2-Streaming's accuracy claim on your own audio
Microsoft says its streaming transcriber ranks first on Artificial Analysis. A 10-clip test on your own audio, with a word error rate script and Sume STT.
Written by Sume