Prometheus histogram buckets for Sume video jobs, up to 20 minutes
Pick histogram buckets for submit-to-terminal time on 30-second Seedance 2.5 and other Sume video jobs, so the 20-minute SDK deadline is the last bucket.

Why default buckets fail on video
Prometheus client libraries ship request-latency buckets that top out around ten seconds. A Sume video job is not a request: the submit returns a job id in the first response and the work runs for minutes. Docs say image jobs often finish inside the 30-second wait, while video, avatar-video and face-swap jobs usually do not.
So record one histogram of seconds from your submit call to the terminal state, and choose buckets that match what you wait for.
Where the top bucket comes from
The TypeScript SDK waits for a job with a 20-minute default timeout (waitForJob), and 20 minutes is the figure the docs call reasonable for video. Make 1200 seconds your last finite bucket: anything above it is a job your client would have stopped waiting for. Seedance 2.5 accepts 4 to 30 seconds of output, so a 30-second clip is the slow end of one model and still fits.
Facts the buckets rest on (read 2026-10-05)
| Fact | Value | Use in the metric |
|---|---|---|
| Sync wait cap | 30 s | Bucket edge at 30 |
| waitForJob default timeout | 20 min | Last finite bucket 1200 |
| waitForRun default timeout | 10 min | Edge at 600 |
| Terminal statuses | completed, failed, canceled | Label status |
The metric
Label by model and terminal status only. Never label by job id, prompt or user: every new label value is a new time series.
import time
from prometheus_client import Histogram
JOB_SECONDS = Histogram(
"sume_job_seconds", "Submit to terminal",
["model", "status"],
buckets=(5, 10, 20, 30, 60, 120, 300, 600, 900, 1200),
)
def observe(model, started, terminal_status):
JOB_SECONDS.labels(model, terminal_status).observe(time.monotonic() - started)
started = time.monotonic()
observe('seedance-2.5', started, 'completed')Alerting on it
Alert on the share of jobs above 600 seconds, not on a single slow job. A client-side timeout does not cancel a Sume job: it keeps running and keeps billing, so a rising tail is a cost signal as well as a latency signal. Store the job id and read it again rather than submitting a second paid request.
Sources
Related posts
More in Developers
- Promise.allSettled for a wave of Sume jobs: keep the partial wins
Promise.all throws away nine good Sume videos when one job fails. Use Promise.allSettled over a wave sized to your plan's concurrency, in 30 lines of Node.
- provider_submission_failed 502: Sume could not start the job, retry
provider_submission_failed is a 502 meaning the job never started. By default it is retryable with retry_after_seconds 30 and next_action retry_later.
- Pydantic v2 models for the Sume /v1/videos poll response
Typed Pydantic v2 models for POST and GET /v1/videos: five status literals, an optional error string, a url list, and a guard that stops a bad status early.
- Python async generator that yields Sume job status until terminal
Write an async generator in stdlib Python that polls /v1/jobs/:id/status, honours next_poll_after_seconds, and stops at a terminal status. No SDK needed.
Written by Sume