Plan Sume STT requests and cost for any recording length
Sume STT takes 10 minutes per request at $0.01 a minute. A Python planner counts requests and the bill for a podcast, a webinar and 500 hours.

A 47-minute 20-second podcast needs 5 STT requests and bills 48 minutes, which is $0.48 at Sume STT 1.0's $0.01 per audio minute. A 90-minute webinar needs 9 requests and $0.90, and a 500-hour archive needs 3,000 requests and $300. The script below computes all three, and you can swap in your own durations.
The two numbers that decide the plan
Sume STT 1.0 charges $0.01 per audio minute, which is the provider list of $0.008 times 1.25 as shown on the API pricing page, and one request takes at most 10 minutes of audio (the docs give 600 seconds as the maximum duration hint). A longer recording has to be cut into pieces before you submit it.
The planner is conservative about rounding. It bills each piece by its whole minutes, because the docs describe the transcript rate as per ceil(minute) of the duration hint when a reservation is made. If a piece settles to fewer minutes, the final charge on the usage dashboard will be lower than the estimate, not higher.
The planner
This is plain Python with no dependencies. It slices the total into pieces of up to 600 seconds, sums the whole minutes in each piece and multiplies by the rate.
Run it as it stands. The three sample jobs are a podcast, a webinar and an archive, and the only thing you need to change for your own audio is the list of durations in seconds. The function returns the request count, the billed minutes and the dollars, so it also drops into a larger budget script.
import math
RATE_PER_MIN = 0.01 # Sume STT 1.0, USD per audio minute
MAX_REQUEST_S = 600 # one request takes at most 10 minutes
def plan(total_seconds):
pieces, left = [], total_seconds
while left > 0:
piece = min(left, MAX_REQUEST_S)
pieces.append(piece)
left -= piece
minutes = sum(math.ceil(p / 60) for p in pieces)
return len(pieces), minutes, minutes * RATE_PER_MIN
jobs = [
("Podcast, 47 min 20 s", 47 * 60 + 20),
("Webinar, 1 h 30 min", 90 * 60),
("Archive, 500 h", 500 * 3600),
]
for label, seconds in jobs:
n, minutes, cost = plan(seconds)
print(f"{label}: {n} requests, {minutes} minutes, ${cost:,.2f}")What it prints
Running the script above on 2026-10-05 printed three lines, one per job.
- Podcast, 47 min 20 s: 5 requests, 48 minutes, $0.48
- Webinar, 1 h 30 min: 9 requests, 90 minutes, $0.90
- Archive, 500 h: 3000 requests, 30000 minutes, $300.00
Reading the result
The podcast row shows the rounding at work: 47 minutes 20 seconds is four full 10-minute pieces plus a 7-minute 20-second piece, which bills as 8 minutes, so the total is 48 minutes, not 47.33. Splitting at exact 10-minute marks is cheapest, because a piece of 9 minutes 1 second bills as 10 minutes while a piece of exactly 9 minutes bills as 9. The difference is a cent per piece, but it is real money across thousands of pieces, and it is the reason to cut on the clock and not on the nearest pause. If you split at silences instead, you may add a few seconds of rounding per piece. On 500 hours, with a 30-second leftover on each of 3,000 pieces, the worst case adds a few dollars to the $300.
The request count is the number to plan your queue around. A workspace has a concurrency limit by plan, 4 on Pro, 8 on Startup and 20 on Scale per the pricing page, so 3,000 requests are worked through in waves, not at once. If every request counted against a concurrency of 4, that would be 750 waves of four jobs, so check how your workspace limit applies to this route. Plan for the wall-clock time as well as the money, and keep each piece's request id so a failed piece can be re-submitted without redoing the rest.
Limits of the estimate
The script prices speech-to-text only. If you transcribe through video inspect instead, the docs say the transcript rate is added to the inspect reservation, which also carries a compute ceiling, so that route costs more than the figures here. For a standalone recording, use the STT planner above. For a clip you will inspect anyway, read the video inspect docs and add the two. Before a large run, price one real file with the planner, run just that file, and compare the usage line with the estimate.
Sources
Related posts
More in Pricing
- A podcast hour cut into six 1-minute clips costs $2.52 on Sume
Transcribe 60 minutes, trim six clips, caption each and render six one-minute timelines: $2.52 on Sume, line by line from published per-job prices.
- Podcast shorts in 3 markets: 20 clips a month, $12.80 on Sume
Four episodes a month with five clips each, captioned for three markets, is 20 clips and 60 caption jobs: $12.80 on Sume at list, with free translation.
- Price a 12-script ad library: MAI-Voice-2.1, Flash and Sume TTS
Twelve ad scripts from 90 to 1,200 characters cost $0.35 on Sume, $0.1386 on MAI-Voice-2.1 and $0.0945 on Flash. The per-job rounding is where they differ.
- Price a 20-prompt video regression suite: 360p, 3 seconds on Sume
A 20-prompt golden set rendered at 360p and 3 seconds on Gemini Omni costs $2.40 per run on Sume. The arithmetic, a Python cost check and the request body.
Written by Sume