Video over 30 minutes: Sume inspect caps at 1800 s
Sume video inspect and video trim reject sources over 1800 seconds. For a longer recording, split it before import, or use audio detach and STT.

Sume's video inspect and video trim both read sources up to 1800 seconds, and video trim returns at most 900 seconds per cut. So video trim cannot split a recording longer than 1800 seconds: it reports a source over the cap as source_duration_exceeded. Split the file before you import it, with your own tool such as ffmpeg, into pieces under the cap, then inspect each piece.
Sume's inspect is a probe, stills and transcript tool, not a video model, so it does not read a long recording in one pass.
What are the numbers?
The caps differ by step, so plan around the smallest.
| Step | Cap | Error when exceeded |
|---|---|---|
| Video inspect source | 1800 s | Source over the cap is refused |
| Video inspect stills per call | 24 | n/a |
| Video trim source | 1800 s | source_duration_exceeded |
| Video trim output | 0.2 s to 900 s | video_trim_range_empty if over 900 s |
How do I cut a recording that is under the cap?
For a source of 1800 seconds or less, call video trim with start and either end or duration, never both. Use precision exact for a frame-accurate cut; keyframe copies the stream and may start up to one GOP early, in which case read actual_start_seconds from the result.
curl -X POST https://api.sume.com/v1/video-trim \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: trim-part-1" \
-d '{"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4", "start": 0, "duration": 600}'What does the cut cost?
Video trim is $0.02 per job per the docs; there is no provider inference, only worker ffmpeg. Splitting outside Sume costs nothing on Sume.
How do I stitch the findings?
Add each piece's start offset to the timestamps in its transcript so they refer to the original recording.
How should I choose the cut points?
Cut where the content changes, not on a fixed grid, when you can. A silence or a chapter marker is a clean place to cut, and a transcript from a first pass of the whole audio can show you where they fall. If you need a simple grid, use pieces of 600 seconds with a few seconds of overlap, so a sentence at the boundary appears whole in at least one piece.
Keep the overlap in mind when you merge results. A claim found in the overlap will appear in two pieces, so de-duplicate by the original-timeline timestamp rather than by the text.
Does the split change the cost?
Splitting outside Sume adds no trim jobs. Each inspect with a transcript is billed by the audio minute, and probe and stills are unbilled. Run the cost through the catalog before you start a large batch.
Where do I split a recording over 1800 seconds?
Before import. Cut the file locally into pieces under 1800 seconds, import each piece, and then inspect or trim it. Keep the original file, so you can re-cut with different boundaries if the first split was poor.
Sources
More in Developers
- Long video understanding: read the transcript, then pull frames
Google's changelog says Gemini fetches transcripts, frames or audio on demand, with 88% fewer tokens. A transcript-first, frames-second flow on Sume inspect.
- Reduce video vendor lock-in: swap the model field, keep the pipeline
Keep the video model id in config, read each model's limits from the catalog, and your pipeline survives a vendor change. A script that lists limits.
- Review voice agent call recordings: STT word timings on Sume
After you ship a Gemini Live or other voice agent, transcribe the recordings with Sume STT: word timings, sentence segments, a 10-minute cap per request.
- Voice API deadlines, October 2026 to February 2027
A calendar of voice and transcription API changes from vendor pages: Gemini TTS price rise, OpenAI transcription shutdown, and the xAI voice alias move.
Written by Sume