Cut a long talk into short vertical clips with the Sume API
Turn one long recording into short clips on Sume: video-inspect for a sentence transcript, video-trim per range, video-captions to burn the words. Rates inside.

To turn one long recording into short clips with Sume, import it, run video-inspect with transcribe: true and sentence segmentation, choose ranges from the segments, cut each range with video-trim, then burn captions with video-captions. The source can be at most 1800 seconds (30 minutes) for inspect and trim. Choosing which sentences make a good clip is your step, or your agent's.
The four calls
| Step | Endpoint | Limit or rate |
|---|---|---|
| Import the file | POST /v1/media-imports | The inspect and trim APIs accept only media.sume.com URLs |
| Transcript | POST /v1/video-inspect with transcribe: true | $0.01 per audio minute, plus compute; source up to 1800 s |
| Cut a range | POST /v1/video-trim | $0.02 per job; output 0.2 to 900 s |
| Burn captions | POST /v1/video-captions | $0.20 per job for a video up to 60 s |
Step by step
The whole flow is asynchronous after the import. Writes default to async for trim and captions, while inspect defaults to a sync wait of up to 30 seconds and returns a queued job if it takes longer. Treat every response as possibly a 202 and poll the job envelope, and you will not need two code paths.
- Inspect with
frames: falseif you only need the transcript, andsegmentation: { mode: "sentence" }for caption-line shapedsegments[]with no gaps. Sendlanguage_codeas a hint, or leave it for auto-detect. A silent clip answersinspect_source_has_no_audio. - Choose ranges from the segments, so a cut starts at a sentence start and ends at a sentence end.
- Trim with
startplus eitherendorduration.precision: exactis the default and frame-accurate.keyframeis a stream copy that can start up to one group of pictures early, so re-base your times onactual_start_seconds. - Send the trimmed
video_urlto captions. Each write needs anIdempotency-Key; poll the job at/v1/jobs/{id}/status.
What it adds up to
For a 30-minute talk cut into eight clips, the listed rates are 30 minutes of transcript at $0.01, eight trims at $0.02 and eight caption jobs at $0.20: $0.30 + $0.16 + $1.60 = $2.06. The inspect job also reserves its compute ceiling, so that figure is the sum of the per-unit rates and not a quote. Read the live prices from GET /v1/catalog.
The vertical part
video-trim keeps the source's frame unless you ask for an output of width, height and fps. If you need a vertical frame that reframes the shot, timeline-1.0 composes slots on video[] with fit set to cover, contain, stretch or blur, and renders 1080 by 1920 by default at $0.10 per started output minute. Run POST /v1/timeline-1.0/plan first; it is unbilled and returns estimated_cost_usd_micros.
The hosted MCP server exposes the same steps as video_trim and the other media tools, each followed by jobs_wait and jobs_result, which suits an agent that picks the ranges itself. Writes there need an idempotency_key, so build one from the source id and the range, and a retry of the same cut will not create a second file.
Sources
Related posts
More in Use cases
- Macro liquid pour video prompt for Gemini Omni and its price
A prompt recipe for a macro pour shot (oil, serum, honey) with Gemini Omni Flash on Sume: camera words, sound line, 5-second cost at 360p, 720p and 1080p.
- A 2.5% word error rate is 25 wrong words per 1,000: plan the proofread
If a 2.5% WER held on your audio, a 1,000-word transcript would still carry about 25 wrong words. How to proofread and burn captions on Sume with script_text.
- MAI-Voice-2.1 pt-BR: six voices, four male, for ad casting
MAI-Voice-2.1 lists six pt-BR voices: four male, two female. Four list excited, two have whispering and shouting. The cast table, and Sume's pt language tag.
- MAI-Voice-2.1 Mandarin: six zh-CN voices, from 6 to 19 styles
Microsoft lists six zh-CN voices for MAI-Voice-2.1, with style lists from 6 to 19 items. Which to audition for an ad, and what to set on Sume for Chinese.
Written by Sume