Cost to transcribe a 3-minute Reel: $0.01 per audio minute
Video inspect with transcribe adds STT at $0.01 per audio minute to the clip's compute. Hints, sentence segments, silent-clip errors, and a check-first order.

Transcribing a 3-minute Reel with Sume costs the speech-to-text rate of $0.01 per audio minute, so the speech part of the job is about $0.03, plus the clip's own compute. Video inspect reserves its Modal compute ceiling and captures the actual container seconds at list price times 1.25 plus the platform fee, never above the hold, and the transcript adds its per-minute rate to that reservation.
Why this fits Reels
Instagram says it transcribes spoken audio into captions automatically (read 2026-10-03) and lets you review and style them. A transcript you have yourself is useful for something else: a post caption, a script for the next version, or sentence-level cues for a design you control.
Options
Set transcribe: true on POST /v1/video-inspect. The extra fields work only with it; sending them without it returns 400 video_inspect_transcribe_required.
| Field | Meaning | Limit |
|---|---|---|
| language_code | STT hint, such as en | Omit to auto-detect |
| segmentation.mode | sentence returns gapless segments[] | silence_split_seconds 0.2 to 3 |
| duration_seconds | Length hint for the reservation | Max 600 s; omit and 1 minute is reserved |
| frames | false gives probe only | Default is 8 stills |
Check first, then transcribe
A silent clip fails with inspect_source_has_no_audio. A probe with frames: false tells you probe.has_audio before you pay for the speech part. Both submits answer 202 with a job, so poll the status URL, then read probe.has_audio from GET /v1/video-inspect/:id before sending the second call.
import os, requests
API = "https://api.sume.com/v1/video-inspect"
h = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
url = os.environ["REEL_URL"]
probe = requests.post(API, headers={**h, "Idempotency-Key": "probe-001"},
json={"video_url": url, "frames": False}).json()
print(probe)
r = requests.post(API, headers={**h, "Idempotency-Key": "tx-001"}, json={
"video_url": url,
"frames": False,
"transcribe": True,
"duration_seconds": 180,
"segmentation": {"mode": "sentence"},
})
print(r.status_code)What Sume does not do
The transcript is not a caption file. To burn text onto the picture, pass authored cues or let video captions run its own recognition; the docs say SRT uploads are not supported. Read any transcript before you reuse it.
Sources
Related posts
More in Pricing
- How Sume rounds video cost to credits: why 5 s is 32, not 31.25
Sume's estimator rounds the billable amount up to whole cents (credits). A 5 s H3 Max 480p clip is $0.3125 and reserves 32 credits; ten of them reserve 320.
- Ideogram 4.5 for 100 SKUs: low, medium, high totals for a mixed plan
Fal lists Ideogram 4.5 at $0.03, $0.06 and $0.22 per image by quality. The totals for 100 SKUs under five plans, and what Sume's docs say about price.
- Ideogram 4.5 on Sume: 1K vs 2K cost the same, quality sets price
Sume's docs list Ideogram 4.5 at $0.03, $0.06 or $0.22 per image by quality, whatever the size. Resolution is 1K or 2K, five references, default medium.
- Ideogram 4.5 four images in one request: what Sume reserves
A four-image Ideogram 4.5 request reserves $0.15 at low, $0.30 at medium and $1.10 at high on Sume, and a 402 means the balance can't cover that hold.
Written by Sume