Prometheus histogram buckets for Sume video jobs, up to 20 minutes

Pick histogram buckets for submit-to-terminal time on 30-second Seedance 2.5 and other Sume video jobs, so the 20-minute SDK deadline is the last bucket.

4 min readSume
All posts

Why default buckets fail on video

Prometheus client libraries ship request-latency buckets that top out around ten seconds. A Sume video job is not a request: the submit returns a job id in the first response and the work runs for minutes. Docs say image jobs often finish inside the 30-second wait, while video, avatar-video and face-swap jobs usually do not.

So record one histogram of seconds from your submit call to the terminal state, and choose buckets that match what you wait for.

Where the top bucket comes from

The TypeScript SDK waits for a job with a 20-minute default timeout (waitForJob), and 20 minutes is the figure the docs call reasonable for video. Make 1200 seconds your last finite bucket: anything above it is a job your client would have stopped waiting for. Seedance 2.5 accepts 4 to 30 seconds of output, so a 30-second clip is the slow end of one model and still fits.

Facts the buckets rest on (read 2026-10-05)

From the Sume jobs and SDK docs, read 2026-10-05
FactValueUse in the metric
Sync wait cap30 sBucket edge at 30
waitForJob default timeout20 minLast finite bucket 1200
waitForRun default timeout10 minEdge at 600
Terminal statusescompleted, failed, canceledLabel status

The metric

Label by model and terminal status only. Never label by job id, prompt or user: every new label value is a new time series.

import time
from prometheus_client import Histogram

JOB_SECONDS = Histogram(
    "sume_job_seconds", "Submit to terminal",
    ["model", "status"],
    buckets=(5, 10, 20, 30, 60, 120, 300, 600, 900, 1200),
)

def observe(model, started, terminal_status):
    JOB_SECONDS.labels(model, terminal_status).observe(time.monotonic() - started)

started = time.monotonic()
observe('seedance-2.5', started, 'completed')

Alerting on it

Alert on the share of jobs above 600 seconds, not on a single slow job. A client-side timeout does not cancel a Sume job: it keeps running and keeps billing, so a rising tail is a cost signal as well as a latency signal. Store the job id and read it again rather than submitting a second paid request.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume