Sume queue position and ETA: none exists, poll generation_limits

Sume shows queue counts and remaining capacity, not a per-job position or ETA. Read generation_limits and keep new work inside the headroom formula.

4 min readSume
All posts

Sume does not give a queued job a position in line or an ETA. It shows queue counts and the remaining accepted capacity in generation_limits, and the documented client behavior is to store the job id, poll with backoff, and keep new work inside the headroom those counts imply.

What you can read

Generation submit responses include a generation_limits snapshot when it can be computed. The counts can change right after the response, since workers claim jobs and other clients submit work.

generation_limits fields for progress estimates, from Generation admission (read 2026-10-05)
FieldUse
concurrency_limitThe effective processing cap of the workspace.
active_generation_jobsJobs in processing now.
queued_generation_jobsJobs waiting in queued.
queue_capacity_remainingQueued budget plus idle processing seats before queue_full.
wave_size_hintA submission-wave hint only; never a concurrency figure.

The one honest estimate

Your headroom for new in-flight work is max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs), limited to queue_capacity_remaining. At zero, wait and refresh before submitting more. That tells you whether a new job will start soon; it does not tell you when.

A job's wait depends on how long the jobs ahead of it take, which varies by model and request, so any ETA you print would be a guess. Show a user the states queued and processing and the counts, not a countdown.

def headroom(gl):
    room = max(0, gl["concurrency_limit"]
               - gl["active_generation_jobs"]
               - gl["queued_generation_jobs"])
    return min(room, gl["queue_capacity_remaining"])

snapshot = {"concurrency_limit": 4, "active_generation_jobs": 3,
            "queued_generation_jobs": 0, "queue_capacity_remaining": 21}
print(headroom(snapshot))  # 1

What to tell users

A queued job is normal, not a failure. Poll the status with exponential backoff until terminal: true. Use the status_url the job returned, not one you build. If you hit 429 queue_full, stop adding work and retry with the same idempotency key after something finishes.

The sync and subscribe modes wait for at most 30 seconds, after which you continue by polling the job id.

A worked example

Take a workspace with concurrency_limit of 4 and a queue of 20, so the accepted capacity is 24. If three jobs are processing and none is queued, the formula gives max(0, 4 - 3 - 0), which is 1, and the queue still has room. Submit one more job and wait for the next snapshot rather than sending five at once.

This keeps work inside the processing cap, while the API would still accept more as queued under its own queue-first policy. The difference matters for latency: a job you hold back on the client is a job that does not sit behind others.

If your submit responses do not carry generation_limits, refresh the counts from a fresh response before choosing a batch width instead of guessing.

Limits

Queue expiry and a client-chosen fail-fast queue length are not public options unless the live OpenAPI schema lists them. Do not build on either until it does.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume