Cost to transcribe a 3-minute Reel: $0.01 per audio minute

Video inspect with transcribe adds STT at $0.01 per audio minute to the clip's compute. Hints, sentence segments, silent-clip errors, and a check-first order.

5 min readSume
All posts

Transcribing a 3-minute Reel with Sume costs the speech-to-text rate of $0.01 per audio minute, so the speech part of the job is about $0.03, plus the clip's own compute. Video inspect reserves its Modal compute ceiling and captures the actual container seconds at list price times 1.25 plus the platform fee, never above the hold, and the transcript adds its per-minute rate to that reservation.

Why this fits Reels

Instagram says it transcribes spoken audio into captions automatically (read 2026-10-03) and lets you review and style them. A transcript you have yourself is useful for something else: a post caption, a script for the next version, or sentence-level cues for a design you control.

Options

Set transcribe: true on POST /v1/video-inspect. The extra fields work only with it; sending them without it returns 400 video_inspect_transcribe_required.

Transcript options on video inspect (read 2026-10-03)
FieldMeaningLimit
language_codeSTT hint, such as enOmit to auto-detect
segmentation.modesentence returns gapless segments[]silence_split_seconds 0.2 to 3
duration_secondsLength hint for the reservationMax 600 s; omit and 1 minute is reserved
framesfalse gives probe onlyDefault is 8 stills

Check first, then transcribe

A silent clip fails with inspect_source_has_no_audio. A probe with frames: false tells you probe.has_audio before you pay for the speech part. Both submits answer 202 with a job, so poll the status URL, then read probe.has_audio from GET /v1/video-inspect/:id before sending the second call.

import os, requests

API = "https://api.sume.com/v1/video-inspect"
h = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
url = os.environ["REEL_URL"]

probe = requests.post(API, headers={**h, "Idempotency-Key": "probe-001"},
                      json={"video_url": url, "frames": False}).json()
print(probe)

r = requests.post(API, headers={**h, "Idempotency-Key": "tx-001"}, json={
    "video_url": url,
    "frames": False,
    "transcribe": True,
    "duration_seconds": 180,
    "segmentation": {"mode": "sentence"},
})
print(r.status_code)

What Sume does not do

The transcript is not a caption file. To burn text onto the picture, pass authored cues or let video captions run its own recognition; the docs say SRT uploads are not supported. Read any transcript before you reuse it.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume