Transcribe a 30-minute recording in three 600-second STT jobs: $0.33
Sume STT takes at most 600 seconds per job. Cut a 30-minute video into three 10-minute mp3 pieces with audio detach and transcribe each: $0.33 total.

The plan and the price
Cut the video into three 600-second ranges with POST /v1/audio-detach, then run POST /v1/stt-1.0/transcribe on each piece. Three detach jobs at $0.01 each plus 30 minutes of STT at $0.01 per minute is $0.03 + $0.30 = $0.33.
Two limits force the split. Sume STT accepts duration_seconds from 1 to 600, so one job covers at most 10 minutes. Audio detach reads a source of up to 1,800 seconds, which is exactly 30 minutes, and writes at most 900 seconds per output.
Why mp3 for the pieces
Sume's STT size guide puts a 10 MB ceiling on the audio file. A 600-second mono wav at 16 kHz and 16 bits is 32,000 bytes a second, or 19.2 MB, which is too big. Detach's mp3 option is 128 kbps, or 16,000 bytes a second, so 600 seconds is 9.6 MB and fits. Ask for format: "mp3" and channels: "mono".
| Piece | Range (seconds) | Detach cost | STT cost (10 min) |
|---|---|---|---|
| 1 | 0 to 600 | $0.01 | $0.10 |
| 2 | 600 to 1200 | $0.01 | $0.10 |
| 3 | 1200 to 1800 | $0.01 | $0.10 |
| Total | 0 to 1800 | $0.03 | $0.30 |
Code
The video must already be a media.sume.com URL; import it first with POST /v1/media-imports. The script below submits, polls with backoff and never resubmits a paid job. Set VIDEO to your imported URL.
import asyncio, os, httpx
API = "https://api.sume.com"
VIDEO = "https://media.sume.com/artifacts/artf_demo/call.mp4"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
async def run(c, path, body, key):
r = await c.post(API + path, json=body, headers={**H, "Idempotency-Key": key})
r.raise_for_status()
job, delay = r.json()["request_id"], 2
while True:
s = (await c.get(f"{API}/v1/jobs/{job}/status", headers=H)).json()
if s["status"] in ("failed", "canceled"): raise RuntimeError(s)
if s.get("result_ready"): break
await asyncio.sleep(delay); delay = min(delay * 2, 15)
return (await c.get(f"{API}/v1/jobs/{job}/result", headers=H)).json()
async def main():
async with httpx.AsyncClient(timeout=60) as c:
for i in range(3):
rng = {"start": i * 600, "end": (i + 1) * 600}
d = await run(c, "/v1/audio-detach", {"video_url": VIDEO, "format": "mp3", "channels": "mono", "range": rng}, f"call-d{i}")
t = await run(c, "/v1/stt-1.0/transcribe", {"audio_url": d["audio_url"], "duration_seconds": 600}, f"call-t{i}")
print(i, t["text"][:80])
asyncio.run(main())Stitching the transcript
Each piece returns words[] with times counted from the start of that piece. To get one timeline, add 600 seconds to every word time in piece 2 and 1,200 seconds to piece 3. A word may be cut at a boundary; if that matters, shift the ranges to a silent point or overlap them by a couple of seconds and drop the repeated words.
- Pass
duration_secondsso the reservation matches the piece; omitting it reserves one minute. - If a piece is shorter than 600 seconds, pass its real length.
- Poll the status URL; do not resubmit a paid job that looks slow.
- Keep the three job ids so you can fetch results again.
What can go wrong
Most failures are caught before you pay for the second step. Detach refuses a source that is not on the Sume media host (unsupported_media_source), a source longer than 1,800 seconds (source_duration_exceeded) and a video with no audio track (detach_source_has_no_audio). A range that ends before it starts, or runs longer than 900 seconds, returns audio_detach_range_empty.
If a recording is longer than 30 minutes, cut it first with a video trim, or detach it in 1,800-second slices, and then apply the same three-piece plan to each. The arithmetic stays linear: every extra 10 minutes is one more detach at $0.01 and one more STT job at $0.10, so $0.11.
| Recording length | Pieces of 600 s | Total cost |
|---|---|---|
| 30 minutes | 3 | $0.33 |
| 60 minutes | 6 | $0.66 |
| 90 minutes | 9 | $0.99 |
Sources
Related posts
More in Developers
- Transcribe a clip when you do not know the language: STT auto-detect
Omit language_code on Sume speech-to-text and the job auto-detects the language. What the result returns, the 10-minute cap, and the $0.01 a minute rate.
- Transcribe audio from a private bucket: Sume STT needs a public URL
Sume STT takes a public HTTPS audio_url. If your files sit in a private bucket, here is how to hand them over without opening the whole bucket.
- 150-language subtitles: which scripts Sume captions document
A translation model can output 150 languages, but Sume documents Latin and Hangul caption styles. Test other scripts on a short clip before a batch.
- Trigger.dev Node 21 warning: which Node runs the Sume SDK
Trigger.dev v4.6.1 added Node.js 21 deprecation warnings. The Sume TypeScript SDK needs Node 18 or later, so tasks on Node 22 or newer are fine.
Written by Sume