Plan Sume STT requests and cost for any recording length

Sume STT takes 10 minutes per request at $0.01 a minute. A Python planner counts requests and the bill for a podcast, a webinar and 500 hours.

5 min readSume
All posts

A 47-minute 20-second podcast needs 5 STT requests and bills 48 minutes, which is $0.48 at Sume STT 1.0's $0.01 per audio minute. A 90-minute webinar needs 9 requests and $0.90, and a 500-hour archive needs 3,000 requests and $300. The script below computes all three, and you can swap in your own durations.

The two numbers that decide the plan

Sume STT 1.0 charges $0.01 per audio minute, which is the provider list of $0.008 times 1.25 as shown on the API pricing page, and one request takes at most 10 minutes of audio (the docs give 600 seconds as the maximum duration hint). A longer recording has to be cut into pieces before you submit it.

The planner is conservative about rounding. It bills each piece by its whole minutes, because the docs describe the transcript rate as per ceil(minute) of the duration hint when a reservation is made. If a piece settles to fewer minutes, the final charge on the usage dashboard will be lower than the estimate, not higher.

The planner

This is plain Python with no dependencies. It slices the total into pieces of up to 600 seconds, sums the whole minutes in each piece and multiplies by the rate.

Run it as it stands. The three sample jobs are a podcast, a webinar and an archive, and the only thing you need to change for your own audio is the list of durations in seconds. The function returns the request count, the billed minutes and the dollars, so it also drops into a larger budget script.

import math

RATE_PER_MIN = 0.01   # Sume STT 1.0, USD per audio minute
MAX_REQUEST_S = 600   # one request takes at most 10 minutes

def plan(total_seconds):
    pieces, left = [], total_seconds
    while left > 0:
        piece = min(left, MAX_REQUEST_S)
        pieces.append(piece)
        left -= piece
    minutes = sum(math.ceil(p / 60) for p in pieces)
    return len(pieces), minutes, minutes * RATE_PER_MIN

jobs = [
    ("Podcast, 47 min 20 s", 47 * 60 + 20),
    ("Webinar, 1 h 30 min", 90 * 60),
    ("Archive, 500 h", 500 * 3600),
]
for label, seconds in jobs:
    n, minutes, cost = plan(seconds)
    print(f"{label}: {n} requests, {minutes} minutes, ${cost:,.2f}")

What it prints

Running the script above on 2026-10-05 printed three lines, one per job.

  • Podcast, 47 min 20 s: 5 requests, 48 minutes, $0.48
  • Webinar, 1 h 30 min: 9 requests, 90 minutes, $0.90
  • Archive, 500 h: 3000 requests, 30000 minutes, $300.00

Reading the result

The podcast row shows the rounding at work: 47 minutes 20 seconds is four full 10-minute pieces plus a 7-minute 20-second piece, which bills as 8 minutes, so the total is 48 minutes, not 47.33. Splitting at exact 10-minute marks is cheapest, because a piece of 9 minutes 1 second bills as 10 minutes while a piece of exactly 9 minutes bills as 9. The difference is a cent per piece, but it is real money across thousands of pieces, and it is the reason to cut on the clock and not on the nearest pause. If you split at silences instead, you may add a few seconds of rounding per piece. On 500 hours, with a 30-second leftover on each of 3,000 pieces, the worst case adds a few dollars to the $300.

The request count is the number to plan your queue around. A workspace has a concurrency limit by plan, 4 on Pro, 8 on Startup and 20 on Scale per the pricing page, so 3,000 requests are worked through in waves, not at once. If every request counted against a concurrency of 4, that would be 750 waves of four jobs, so check how your workspace limit applies to this route. Plan for the wall-clock time as well as the money, and keep each piece's request id so a failed piece can be re-submitted without redoing the rest.

Limits of the estimate

The script prices speech-to-text only. If you transcribe through video inspect instead, the docs say the transcript rate is added to the inspect reservation, which also carries a compute ceiling, so that route costs more than the figures here. For a standalone recording, use the STT planner above. For a clip you will inspect anyway, read the video inspect docs and add the two. Before a large run, price one real file with the planner, run just that file, and compare the usage line with the estimate.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume