What one captioned Shorts episode costs on Sume, itemized
Add the documented Sume rates for one Shorts episode: render, trim, captions, transcript. A Python estimator refuses captions past the 60-second quote.

The line items
One captioned Shorts episode on Sume is a sum of at most four documented charges: a Timeline render, an optional trim, the captions job and, if you read a transcript to choose the range, a video inspect transcription. Each line is a rate the Sume docs print on that endpoint's page. They are not a promise: every page says to confirm in GET /v1/catalog. YouTube's help page says a Short can run up to three minutes, but the captions price is quoted for videos up to 60 seconds, so a captioned episode in this estimate is at most a minute.
We list no total for a season because the stored posts already do that; this post gives you the line items for one episode and an estimator that tells you when you have left the quoted range.
Rates and what they bill on
Billing units differ per row: per job, per whole output minute and per audio minute. That difference is where estimates go wrong.
| Step | Rate | Billed per | Condition |
|---|---|---|---|
| Timeline render | $0.10 | Whole output minute | A 45 s episode bills one minute |
| Video trim | $0.02 | Job | Optional; only if you cut a longer file |
| Video captions | $0.20 | Job | Quoted for videos up to 60 s |
| Video inspect transcription | $0.01 | Audio minute | Optional; to choose a range from text |
| Timeline plan | Unbilled | Call | Returns billable_minutes before you render |
An estimator you can run
The function below uses those four numbers and raises when you ask for captions on a clip over 60 seconds, since the docs quote no price there. For a 45-second episode with captions and no trim it prints 0.3; for the same episode cut from a five-minute file and transcribed first it prints 0.37. The third call prints 0.1.
import math
TRIM, CAPTIONS, INSPECT_MIN, TIMELINE_MIN = 0.02, 0.20, 0.01, 0.10 # documented rates
def episode_cost(seconds, trimmed_from=None, captioned=True, transcribed=False):
"""Documented-rate estimate in USD for one episode. Confirm rates in GET /v1/catalog."""
total = math.ceil(seconds / 60) * TIMELINE_MIN
if trimmed_from:
total += TRIM
if captioned:
if seconds > 60:
raise ValueError("captions are priced for videos up to 60 s")
total += CAPTIONS
if transcribed:
total += math.ceil((trimmed_from or seconds) / 60) * INSPECT_MIN
return round(total, 2)
print(episode_cost(45)) # plain 45 s episode
print(episode_cost(45, trimmed_from=300, transcribed=True)) # cut from 5 min, transcribed
print(episode_cost(58, captioned=False))What the estimate leaves out
It leaves out the cost of making the clips: image and video generation are separate endpoints with their own prices, and the catalog is where you read them. It leaves out any re-render you pay for after a warning, because a re-render is a new job at the same rate. And it leaves out the time you spend; a render is a job, not an immediate answer.
The estimator rounds the render up to whole minutes, which is how the docs state the unit. A 61-second episode costs two minutes, so trimming one second off a 61-second cut saves $0.10, which is half of the caption price. That is a useful habit for a series: aim for 59 seconds.
Use it with the plan
The unbilled plan is the real authority on the render line, because it knows your actual duration and segment count. Use the estimator for the parts the plan does not cover, such as trim and captions, and compare its render line against billable_minutes from the plan. If they disagree, trust the plan and look at your code.
Worked through by hand, the three calls above come out like this. A 45-second episode that needs no trim is one render minute at $0.10 plus the $0.20 captions job, so $0.30. The same episode cut from a five-minute recording adds the $0.02 trim, and if you read the whole five minutes of transcript first, five audio minutes at $0.01 add $0.05, so $0.37. A 58-second episode with no captions is just the render minute, $0.10.
Two habits make the estimate more useful. First, keep the rates in one place in your code so a catalog change is a one-line edit. Second, print the breakdown, not just the total, so a reviewer can see which step a surprise came from. A total of $0.37 tells nobody anything; a trim of $0.02 plus a transcript of $0.05 does.
If you also keep a ledger of jobs, add the estimator's output as a column next to the actual billed figure. Over a season, the gap between the two tells you whether your assumptions about trims and transcripts are right. A gap that grows is a sign that your process is using steps the estimate does not count.
For a three-minute Short the picture changes. The render line rises to three minutes, and the captions line has no quoted price above 60 seconds, so the estimator raises on purpose. For long episodes, read the live catalog and OpenAPI before you commit, and do not assume the 60-second rate carries over.
Sources
Related posts
More in Pricing
- What one Ideogram 4.5 image costs in voice minutes and video seconds
One medium Ideogram 4.5 image costs as much as 1.75 minutes of Sume speech or 7.5 transcript minutes. An exchange-rate table from images to audio and video.
- How Sume pricing works: plans, one wallet, published model rates
Sume plans set access and concurrency. Usage draws from one prepaid wallet at each model's published USD rate, for generation, the Agent, Formats, and the API.
- AI avatar video API pricing: cost per second and per minute
Sume bills AI avatar video per second by quality tier, with separate rates when you send a product image. Per-minute costs for standard, plus, and max.
- AI video generation cost per video: what one Sume API run cost
To see what one Sume run or video job cost, call GET /v1/usage with run_id or job_id and read debited_usd: the wallet deduction, agent turns included.
Written by Sume