Trim and caption a 30-second AI clip: two fixed prices, $0.22
Trim is $0.02 per job and captions are $0.20 per accepted job for videos up to 60 s, so a 30-second AI clip costs $0.22 after the render.

Trimming and captioning one AI clip through the Sume API adds two fixed charges: $0.02 for the trim job and $0.20 for the caption job, so $0.22 on top of the generation itself. The caption price covers a video of up to 60 seconds, so a 30-second clip pays the same $0.20 as a 55-second one. Both numbers come from the docs' current fixed estimates; read GET /v1/catalog for the live price before you budget a batch.
The generation line is the part that scales with length. Seedance 2.5 is described by ByteDance as generating up to 30 seconds per pass (read 2026-10-05), and Sume's catalog lists it at 4 to 30 seconds. Sume bills the video models list price times 1.25 for each output second, so use the catalog rather than a number from a blog post.
The post-generation steps
Neither step calls a model for inference. The trim docs state there is no provider inference and only the worker's ffmpeg runs, which is why the price is a flat per-job number.
| Step | Price | What the price covers | Limit that matters |
|---|---|---|---|
POST /v1/video-trim | $0.02 per job | One cut of one clip into a new MP4 | Source up to 1800 s, output 0.2 to 900 s |
POST /v1/video-captions | $0.20 per accepted job | Burning captions onto a video of up to 60 s | Public HTTPS video_url |
Caption restyle with source_caption_id | $0.20 (a restyle is still a render) | Same video, new style, no second transcription | Needs the earlier caption job's id |
Where the numbers can mislead you
A caption job reserves and captures $0.20 when it is accepted, and a clip over 60 seconds is outside the stated price. Keep caption sources at 60 seconds or shorter, or read the live price for longer ones. The trim output cap of 900 seconds is far above a 30-second clip, so it will not bite here.
The trim step is only worth paying for when you need a different range. If the 30-second render is already what you want to caption, skip trim and send the render's URL straight to captions.
Worked example for a batch
For 50 clips that each need a trim and a caption burn, the post-generation spend is 50 x ($0.02 + $0.20) = $11.00. Generation is separate and comes from the model's per-second rate multiplied by clip length. If you also restyle each caption once with source_caption_id, add another 50 x $0.20 = $10.00, because a restyle is billed as a render.
Paid jobs reserve the estimate when they are accepted and capture on success. If a submit fails with 402 insufficient_credits, Sume could not reserve the estimate from the workspace balance, and no provider work started. Check the balance before a large wave rather than discovering it on item 31.
Where the $0.22 does not apply
The two fixed prices cover the trim and caption steps only. The render that produces the 30-second clip is priced by its own model, and the Video Router catalog is the place to read the model you pick. Add the render price to the $0.22 to get the pipeline total, and read the credit amount from the job rather than from your own table, since prices can change.
Trim is priced per job, not per second, and the source can be up to 1800 seconds with an output of 0.2 to 900 seconds. Captions are priced per accepted job for videos up to 60 seconds. A 30-second clip sits well inside both ranges, so the length does not move either price.
A restyle through source_caption_id is still a render and still costs $0.20, but it skips speech-to-text because the earlier word timings are reused. Budget one caption charge per style you want to compare.
- Trim: $0.02 per job.
- Captions: $0.20 per accepted job, videos up to 60 s.
- Render: the model's own price, read from the catalog.
- A failed admission, such as
insufficient_credits, does not start provider work.
Sources
Related posts
More in Developers
- Korean TTS segment text has no spaces, unless a digit is in it
Sume's TTS segment text joins tokens with spaces only if one has a Latin letter or digit; else with nothing. Use segments for timing, your script for text.
- TTS sentence slices: mp3 gives timings only, wav gives audio_urls
Sume TTS segmentation returns sentence timings for any container, but slice audio_urls only with wav or raw. Request shape, the 70 ms rule and when to pick wav.
- Pipes and @{} markers in a Sume TTS transcript: stripped, never spoken
Sume strips || cue breaks, @{...} markers and the display side of <display|spoken> before the voice reads; an empty result returns 400 transcript_no_speech.
- TTS with only an API key: list avatars and pick one with voice ready
You do not need a voice id for Sume TTS. List your avatars, pick one whose voice status is ready, and send its handle as avatar_handle. Python example.
Written by Sume