systemd restart: finish Sume jobs with sume jobs watch, never resubmit

A systemd unit that restarts a finisher and resumes stored Sume job ids with sume jobs watch and download, so a crash never pays for a second job.

5 min readSume
All posts

Write each job_id to disk the moment a submit returns, and let a systemd service that restarts on failure walk that directory. For each id it runs sume jobs watch <job_id>, then sume jobs download <job_id> --output-dir ./out, and removes the id file only after the download succeeds. A crash, a reboot or a systemctl restart then resumes the same jobs, and nothing submits a second paid request.

This is the recovery rule from the CLI jobs docs: if a process restarts or a local timeout occurs, use the job recovery commands and do not submit paid generation again. jobs watch polls until the job is terminal or a timeout occurs, and jobs download writes completed media artifacts to a directory. The helpers also work for Image, Video and Music jobs created through the Developer API.

The unit

The unit runs as a long-lived service and restarts after a failure. The API key comes from an environment file, using the SUME_API_KEY variable that the CLI reads.

[Unit]
Description=Finish pending Sume jobs
After=network-online.target
Wants=network-online.target

[Service]
User=sume
EnvironmentFile=/etc/sume/env
ExecStart=/usr/local/bin/finish-sume-jobs.sh
Restart=on-failure
RestartSec=30

[Install]
WantedBy=multi-user.target

The script

The script exits non-zero if any job fails to finish cleanly, which makes systemd restart it after RestartSec. A job that ended failed or canceled is a final answer: record it and remove the id so it is not retried.

#!/usr/bin/env bash
set -u
: "${SUME_API_KEY:?SUME_API_KEY is empty}"
PENDING=/var/lib/sume/pending
OUT=/var/lib/sume/out
mkdir -p "$PENDING" "$OUT"
rc=0
for f in "$PENDING"/*.id; do
  [ -e "$f" ] || continue
  id=$(cat "$f")
  if sume jobs watch "$id" && sume jobs download "$id" --output-dir "$OUT/$id"; then
    rm -f "$f"
  else
    echo "job $id not finished or not downloadable yet" >&2
    rc=1
  fi
done
exit $rc

Which command handles which situation

Match each situation to the command that handles it. None of them is a new submit.

Recovery commands for an interrupted Sume job (Sume docs, read 2026-10-04)
SituationCommand
Process restarted, job still runningsume jobs watch <job_id>
Need the current state oncesume jobs status <job_id> --agent --json
Job completed, files needed locallysume jobs download <job_id> --output-dir ./out
Want the timeline of a slow jobsume jobs events <job_id> --agent --json
Queued job you no longer wantsume jobs cancel <job_id> --confirm-submit

Caveats

  • The CLI does not ship sume video or sume image generators. Submit those jobs through the HTTP API and use the CLI only to recover them.
  • Cancel works only before generation starts. After that the API answers 409 job_generation_already_started and the job finishes normally.
  • A local timeout does not cancel anything. The job keeps running and keeps billing, so keep the id file until the download succeeds.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume