Transcribe the first 60 seconds of 500 support recordings
500 screen-recorded support sessions: detach seconds 0 to 60 ($5.00), then STT 60 s each ($5.00). Total $10.00 at $0.01 per job and per minute.

Transcribing only the first minute of 500 support recordings costs $10.00 on Sume: $5.00 for 500 audio detach jobs at $0.01 each plus $5.00 for 500 minutes of speech-to-text at $0.01 per audio minute. Audio detach takes a range, so you extract just seconds 0 to 60 and never pay to transcribe the rest.
Cost table
Detach is a flat $0.01 per job (docs: AUDIO_DETACH_PUBLIC_PRICING) and STT is $0.01 per audio minute (docs, Video inspect page). The STT line assumes you send duration_seconds: 60 so the reservation matches the clip.
| Step | Unit price | Units | Cost |
|---|---|---|---|
| Audio detach, range 0-60 s, 16 kHz mono wav | $0.01 per job | 500 | $5.00 |
| STT 1.0, 60 s each | $0.01 per audio minute | 500 minutes | $5.00 |
| Total | $10.00 |
The two calls
The recordings must already be media.sume.com videos in your workspace; import them first with POST /v1/media-imports. Detach requires an Idempotency-Key. 16000 Hz with mono channels is the STT shape the detach docs recommend. Wait for the detach job, read audio_url from its result, then post it to STT as a public HTTPS audio_url.
curl -X POST https://api.sume.com/v1/audio-detach \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: support-0001-detach" \
-d '{"video_url": "https://media.sume.com/artifacts/artf_demo/call.mp4",
"range": {"start": 0, "end": 60},
"channels": "mono", "sample_rate": 16000}'
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: support-0001-stt" \
-d '{"audio_url": "<audio_url from the detach result>", "duration_seconds": 60}'Gotchas
A recording with no audio track fails with detach_source_has_no_audio; screen recordings are the usual culprit, so probe has_audio with video inspect (frames false) before a 500-file run if you are unsure.
If you leave duration_seconds out, STT reserves one minute, which is harmless here but wrong for 10-second clips. The maximum is 600 seconds, so longer calls need more than one range. Word timings come back on every STT job, so you can find the greeting and the issue statement inside the first minute without extra flags.
Sources
Related posts
More in Developers
- Trim 12 Shorts from one recording with asyncio and a semaphore
Cut 12 Shorts from one recording with asyncio and httpx: 12 trims at $0.02 is $0.24. A semaphore of 4 caps in-flight requests; the limit is your choice.
- TTS 1.0 rejects a model field: use the TTS router to pick sonic-3.6
Sume TTS 1.0 does not accept model. To choose an engine such as sonic-3.6 use /v1/tts-router/generate. Price is the same $0.0475 per 1,000 characters.
- Sume TTS pace test: 1,000 characters a minute decides the cap
Sume TTS stops at 20,000 characters or 1,200 seconds. The break-even is 1,000 characters a minute; measure your voice before a long read. Max $0.95 a job.
- Sume TTS speed 0.6 to 1.5: a 14-minute script and the 1,200 s cap
generation_config.speed runs 0.6 to 1.5. A 14-minute read slowed to 0.6 would need about 23 minutes and hit the 1,200-second cap; the price stays per character.
Written by Sume