Transcribe 10,000 short clips: queue capacity and wave size by plan

10,000 five-second clips cost $8.33 on Sume STT. How fast they finish depends on plan concurrency and queue capacity: 6, 24, 48 or 120 accepted jobs.

6 min readSume
All posts

To transcribe 10,000 short clips with Sume STT 1.0, submit one POST /v1/stt-1.0/transcribe job per clip and keep the number of open jobs inside your workspace's accepted-job capacity: 6 on Free, 24 on Pro, 48 on Startup and 120 on Scale and Enterprise. Cost is not the constraint. Ten thousand 5-second clips are 50,000 audio seconds, or 833.3 minutes, so $8.33 at the published $0.01 per audio minute.

What changes with the plan is how many jobs can be in the system at once. The numbers below come from Generation admission (read 2026-10-06).

How many jobs can a workspace hold at once?

Concurrency is a dispatch limit, not a submit limit. A valid job is accepted as queued while queue capacity remains, and workers move it to processing as slots open. Accepted capacity is concurrency_limit + queued_jobs_limit; when it is full, the submit fails with 429 queue_full. The wave_size_hint in the submit response is max(1, floor(queue_capacity_remaining * 0.75)), a submission-wave hint and not a concurrency limit.

Default plan limits, from Generation admission (docs.sume.com), read 2026-10-06. The dashboard Concurrency tab is the source of truth for your workspace.
PlanProcessingQueueAccepted jobsWave hint (empty queue)Waves for 10,000 clips
Free15642,500
Pro4202418556
Startup8404836278
Scale2010012090112

What does the whole batch cost?

STT 1.0 is priced per audio minute and the estimate is prorated by the second from duration_seconds, so short clips are not rounded up to a minute. For comparison, the Microsoft MAI-Transcribe-2-Streaming introductory rate of $0.54 per hour (reported by PYMNTS, read 2026-10-06) would be 13.89 hours, or $7.50, for the same 50,000 seconds. The two bills are within a dollar of each other; the work is in submitting and collecting 10,000 results, not in the price.

Same 50,000 audio seconds. Sume rate from the video inspect docs (docs.sume.com), MAI rate as reported by PYMNTS, both read 2026-10-06.
ServiceRateAudioBill
Sume STT 1.0$0.01 per audio minute833.3 minutes$8.33
MAI-Transcribe-2-Streaming (intro, through 2026-12-31)$0.54 per hour13.89 hours$7.50

How should I pace the submits?

Read generation_limits from each submit response when it is present, stop adding work when queue_capacity_remaining reaches 0, and honor retry-after on a 429. Send an Idempotency-Key per clip so a retry after a timeout never creates a second paid job. Send duration_seconds so each job holds seconds, not the one-minute default.

import os, time, requests

API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

def submit(url, seconds, key):
    r = requests.post(
        f"{API}/v1/stt-1.0/transcribe",
        headers={**H, "Idempotency-Key": key},
        json={"audio_url": url, "duration_seconds": seconds},
        timeout=30,
    )
    if r.status_code == 429:
        time.sleep(float(r.headers.get("retry-after", "5")))
        return None
    r.raise_for_status()
    return r.json()


def run(clips):
    jobs = []
    for url, seconds, key in clips:
        body = None
        while body is None:
            body = submit(url, seconds, key)
        jobs.append(body.get("request_id"))
        limits = body.get("generation_limits") or {}
        if limits.get("queue_capacity_remaining") == 0:
            time.sleep(10)
    return jobs

What to do with the job ids

Store every job id the moment it comes back, and do not submit the same clip again because a local process stopped. Collect results by polling with backoff or by joining 20 ids per MCP wait. Match results to your files with the idempotency key each job echoes, as in match Sume STT results to your file ids.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume