Transcribe the first 60 seconds of 500 support recordings

500 screen-recorded support sessions: detach seconds 0 to 60 ($5.00), then STT 60 s each ($5.00). Total $10.00 at $0.01 per job and per minute.

5 min readSume
All posts

Transcribing only the first minute of 500 support recordings costs $10.00 on Sume: $5.00 for 500 audio detach jobs at $0.01 each plus $5.00 for 500 minutes of speech-to-text at $0.01 per audio minute. Audio detach takes a range, so you extract just seconds 0 to 60 and never pay to transcribe the rest.

Cost table

Detach is a flat $0.01 per job (docs: AUDIO_DETACH_PUBLIC_PRICING) and STT is $0.01 per audio minute (docs, Video inspect page). The STT line assumes you send duration_seconds: 60 so the reservation matches the clip.

500 recordings, first 60 s only (read 2026-10-09)
StepUnit priceUnitsCost
Audio detach, range 0-60 s, 16 kHz mono wav$0.01 per job500$5.00
STT 1.0, 60 s each$0.01 per audio minute500 minutes$5.00
Total$10.00

The two calls

The recordings must already be media.sume.com videos in your workspace; import them first with POST /v1/media-imports. Detach requires an Idempotency-Key. 16000 Hz with mono channels is the STT shape the detach docs recommend. Wait for the detach job, read audio_url from its result, then post it to STT as a public HTTPS audio_url.

curl -X POST https://api.sume.com/v1/audio-detach \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: support-0001-detach" \
  -d '{"video_url": "https://media.sume.com/artifacts/artf_demo/call.mp4",
       "range": {"start": 0, "end": 60},
       "channels": "mono", "sample_rate": 16000}'

curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: support-0001-stt" \
  -d '{"audio_url": "<audio_url from the detach result>", "duration_seconds": 60}'

Gotchas

A recording with no audio track fails with detach_source_has_no_audio; screen recordings are the usual culprit, so probe has_audio with video inspect (frames false) before a 500-file run if you are unsure.

If you leave duration_seconds out, STT reserves one minute, which is harmless here but wrong for 10-second clips. The maximum is 600 seconds, so longer calls need more than one range. Word timings come back on every STT job, so you can find the greeting and the issue statement inside the first minute without extra flags.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume