What one captioned Shorts episode costs on Sume, itemized

Add the documented Sume rates for one Shorts episode: render, trim, captions, transcript. A Python estimator refuses captions past the 60-second quote.

6 min readSume
All posts

The line items

One captioned Shorts episode on Sume is a sum of at most four documented charges: a Timeline render, an optional trim, the captions job and, if you read a transcript to choose the range, a video inspect transcription. Each line is a rate the Sume docs print on that endpoint's page. They are not a promise: every page says to confirm in GET /v1/catalog. YouTube's help page says a Short can run up to three minutes, but the captions price is quoted for videos up to 60 seconds, so a captioned episode in this estimate is at most a minute.

We list no total for a season because the stored posts already do that; this post gives you the line items for one episode and an estimator that tells you when you have left the quoted range.

Rates and what they bill on

Billing units differ per row: per job, per whole output minute and per audio minute. That difference is where estimates go wrong.

Documented rates for the steps of one episode (read 2026-10-06)
StepRateBilled perCondition
Timeline render$0.10Whole output minuteA 45 s episode bills one minute
Video trim$0.02JobOptional; only if you cut a longer file
Video captions$0.20JobQuoted for videos up to 60 s
Video inspect transcription$0.01Audio minuteOptional; to choose a range from text
Timeline planUnbilledCallReturns billable_minutes before you render

An estimator you can run

The function below uses those four numbers and raises when you ask for captions on a clip over 60 seconds, since the docs quote no price there. For a 45-second episode with captions and no trim it prints 0.3; for the same episode cut from a five-minute file and transcribed first it prints 0.37. The third call prints 0.1.

import math

TRIM, CAPTIONS, INSPECT_MIN, TIMELINE_MIN = 0.02, 0.20, 0.01, 0.10  # documented rates

def episode_cost(seconds, trimmed_from=None, captioned=True, transcribed=False):
    """Documented-rate estimate in USD for one episode. Confirm rates in GET /v1/catalog."""
    total = math.ceil(seconds / 60) * TIMELINE_MIN
    if trimmed_from:
        total += TRIM
    if captioned:
        if seconds > 60:
            raise ValueError("captions are priced for videos up to 60 s")
        total += CAPTIONS
    if transcribed:
        total += math.ceil((trimmed_from or seconds) / 60) * INSPECT_MIN
    return round(total, 2)

print(episode_cost(45))                                   # plain 45 s episode
print(episode_cost(45, trimmed_from=300, transcribed=True))  # cut from 5 min, transcribed
print(episode_cost(58, captioned=False))

What the estimate leaves out

It leaves out the cost of making the clips: image and video generation are separate endpoints with their own prices, and the catalog is where you read them. It leaves out any re-render you pay for after a warning, because a re-render is a new job at the same rate. And it leaves out the time you spend; a render is a job, not an immediate answer.

The estimator rounds the render up to whole minutes, which is how the docs state the unit. A 61-second episode costs two minutes, so trimming one second off a 61-second cut saves $0.10, which is half of the caption price. That is a useful habit for a series: aim for 59 seconds.

Use it with the plan

The unbilled plan is the real authority on the render line, because it knows your actual duration and segment count. Use the estimator for the parts the plan does not cover, such as trim and captions, and compare its render line against billable_minutes from the plan. If they disagree, trust the plan and look at your code.

Worked through by hand, the three calls above come out like this. A 45-second episode that needs no trim is one render minute at $0.10 plus the $0.20 captions job, so $0.30. The same episode cut from a five-minute recording adds the $0.02 trim, and if you read the whole five minutes of transcript first, five audio minutes at $0.01 add $0.05, so $0.37. A 58-second episode with no captions is just the render minute, $0.10.

Two habits make the estimate more useful. First, keep the rates in one place in your code so a catalog change is a one-line edit. Second, print the breakdown, not just the total, so a reviewer can see which step a surprise came from. A total of $0.37 tells nobody anything; a trim of $0.02 plus a transcript of $0.05 does.

If you also keep a ledger of jobs, add the estimator's output as a column next to the actual billed figure. Over a season, the gap between the two tells you whether your assumptions about trims and transcripts are right. A gap that grows is a sign that your process is using steps the estimate does not count.

For a three-minute Short the picture changes. The render line rises to three minutes, and the captions line has no quoted price above 60 seconds, so the estimator raises on purpose. For long episodes, read the live catalog and OpenAPI before you commit, and do not assume the 60-second rate carries over.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume