Sume jobs_wait outcome: wait_slice_expired is not a failed job

jobs_wait returns outcome terminal, wait_slice_expired or operator_stopped. Only terminal means the jobs ended; an expired slice says nothing about the jobs.

4 min readSume
All posts

The outcome field on a Sume jobs_wait result has three values: terminal, wait_slice_expired and operator_stopped. Only terminal means the waited-for jobs reached a final state. wait_slice_expired means your polling window closed first and says nothing about the jobs, which keep running and stay billed.

Why `timed_out` was not enough

timed_out: true is a scheduling fact about the caller's window. Agents read it as a verdict. The outcome field was added so the ending is named explicitly, and a timed-out result carries guidance in the same object the model is reading when it decides. The guidance for an expired slice is blunt: any job still queued or processing is billed and will reach a terminal state on its own, so re-issue jobs_wait with the same ids, and proceeding without those ids produces a deliverable missing what they were generating.

jobs_wait outcome values (read 2026-10-05 against the Sume codebase)
outcomeMeaningWhat to do
terminalevery waited-for job reached a terminal stateread each job status, then jobs_result for the successes
wait_slice_expiredthe polling window closed firstre-issue jobs_wait on the same ids
operator_stoppedoperations stopped at least one jobstop polling it; wait on pending ids only

Why the slice is short

One jobs_wait call holds a single request open and sends nothing until it answers, so the edge in front of the API bounds it. Remote HTTP clamps the hold to 55 seconds, and an omitted timeout_seconds defaults to 50 there. A longer ask is accepted and clamped, and the result says so. The clamp comparison shows how client timeouts interact.

Loop on outcome

This loop stops only on a terminal outcome. The wait function is a stand-in so it runs as written.

def wait_all(wait, ids, max_slices=40):
    for _ in range(max_slices):
        res = wait(ids)
        if res['outcome'] == 'terminal':
            return res
        if res['outcome'] == 'operator_stopped':
            ids = res.get('pending_job_ids', [])
            if not ids:
                return res
    raise TimeoutError('still running; tell the user, do not ship partial output')

seq = iter([{'outcome': 'wait_slice_expired'}, {'outcome': 'terminal'}])
print(wait_all(lambda i: next(seq), ['job_a']))

Limits

Slice counts are your budget, not Sume's. If your own deadline arrives first, report the jobs as still running with their ids; do not call them failed.

A real failure pattern

A production run read two consecutive 55-second batches with timed_out: true, plus two 409 reads, as a sign that a B-roll layer was not coming. It assembled a talking-head-only video and marked the run complete. Both clips finished minutes later, billed and unused. The loudest field in the response had been about the caller's polling window; the field that mattered was that the jobs were still in progress and not terminal.

Checklist before you ship

  • Branch on outcome, with timed_out as a hint only.
  • Never assemble a deliverable from a batch that still has pending ids.
  • Keep the full id list and re-issue the same wait.
  • Treat a job as failed only when its own status says failed.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume