Cut a long talk into short vertical clips with the Sume API

Turn one long recording into short clips on Sume: video-inspect for a sentence transcript, video-trim per range, video-captions to burn the words. Rates inside.

6 min readSume
All posts

To turn one long recording into short clips with Sume, import it, run video-inspect with transcribe: true and sentence segmentation, choose ranges from the segments, cut each range with video-trim, then burn captions with video-captions. The source can be at most 1800 seconds (30 minutes) for inspect and trim. Choosing which sentences make a good clip is your step, or your agent's.

The four calls

Long-to-short pipeline (read 2026-10-07)
StepEndpointLimit or rate
Import the filePOST /v1/media-importsThe inspect and trim APIs accept only media.sume.com URLs
TranscriptPOST /v1/video-inspect with transcribe: true$0.01 per audio minute, plus compute; source up to 1800 s
Cut a rangePOST /v1/video-trim$0.02 per job; output 0.2 to 900 s
Burn captionsPOST /v1/video-captions$0.20 per job for a video up to 60 s

Step by step

The whole flow is asynchronous after the import. Writes default to async for trim and captions, while inspect defaults to a sync wait of up to 30 seconds and returns a queued job if it takes longer. Treat every response as possibly a 202 and poll the job envelope, and you will not need two code paths.

  • Inspect with frames: false if you only need the transcript, and segmentation: { mode: "sentence" } for caption-line shaped segments[] with no gaps. Send language_code as a hint, or leave it for auto-detect. A silent clip answers inspect_source_has_no_audio.
  • Choose ranges from the segments, so a cut starts at a sentence start and ends at a sentence end.
  • Trim with start plus either end or duration. precision: exact is the default and frame-accurate. keyframe is a stream copy that can start up to one group of pictures early, so re-base your times on actual_start_seconds.
  • Send the trimmed video_url to captions. Each write needs an Idempotency-Key; poll the job at /v1/jobs/{id}/status.

What it adds up to

For a 30-minute talk cut into eight clips, the listed rates are 30 minutes of transcript at $0.01, eight trims at $0.02 and eight caption jobs at $0.20: $0.30 + $0.16 + $1.60 = $2.06. The inspect job also reserves its compute ceiling, so that figure is the sum of the per-unit rates and not a quote. Read the live prices from GET /v1/catalog.

The vertical part

video-trim keeps the source's frame unless you ask for an output of width, height and fps. If you need a vertical frame that reframes the shot, timeline-1.0 composes slots on video[] with fit set to cover, contain, stretch or blur, and renders 1080 by 1920 by default at $0.10 per started output minute. Run POST /v1/timeline-1.0/plan first; it is unbilled and returns estimated_cost_usd_micros.

The hosted MCP server exposes the same steps as video_trim and the other media tools, each followed by jobs_wait and jobs_result, which suits an agent that picks the ranges itself. Writes there need an idempotency_key, so build one from the source id and the range, and a retry of the same cut will not create a second file.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume