Trim, caption, compose: Sume returns new files, keeping your original
Sume's trim and caption tools return new MP4s and leave the source unchanged. Why that keeps the original AI generation intact, and what each step costs.

Yes: Sume's trim returns a new MP4 and leaves the source unchanged, so the original AI generation stays as it was. The same holds for the other prepare steps, which each make a new output from a source you keep. That lets you cut, caption and assemble for each platform and still have the master to go back to.
The prices and limits below are from the Sume docs for video trim, video captions, timeline and video inspect.
The steps and what they cost
Public rates as listed in the docs.
| Step | Route | Rate | Notes |
|---|---|---|---|
| Trim | POST /v1/video-trim | $0.02 per job | Cut 0.2 to 900 s; source up to 1800 s; new MP4 |
| Captions | POST /v1/video-captions | $0.20 per job | Videos up to 60 s; authored cues or speech-to-text |
| Inspect | POST /v1/video-inspect | $0.01 per audio minute for speech-to-text | Probe and up to 24 stills; source up to 1800 s |
| Timeline render | POST /v1/timeline-1.0/render | $0.10 per ceil output minute | Default output 1080x1920; plan is unbilled |
A safe order
Work from the master, one step at a time, and keep every output with the job id that made it.
- Inspect the master first. The probe tells you duration, size and whether it has audio before you pay for anything else.
- Trim to the length the platform allows.
- Caption the trimmed file. A caption job restyles without re-transcribing when you pass
source_caption_id. - Assemble longer sequences with the timeline render, and use its
/plancall to check a document before you pay for a render.
Why this matters for provenance
Adobe's guidance, covered in this post, is that re-encoding steps can drop Content Credentials, and the delivered-file check shows how to test a final file. Sume's docs list no Content Credentials signing step for outputs, and the exact trim re-encodes the video (libx264, yuv420p), while keyframe precision copies the stream. Do not assume metadata survives either; test the final file you will post.
Keep a ledger
One row per output: master job id, step, output URL, cost, date and the platform it went to. The cost column comes from the poll responses and the rates above, so there is no estimating involved. When a platform asks how a video was made, you can point to the master, not guess at it.
Sources
Related posts
More in Media tools
- Two-voice dialogue audio on Sume: TTS lines joined with audio concat
Build a role-play or interview track by generating one TTS line per turn in each speaker's voice, then joining up to 20 turns into one gapless file for $0.01.
- One lip-sync model per video: Fabric runs at 25 fps, measure H3 Max
Mixing Fabric and MiniMax H3 Max lip-sync clips in one Sume video risks a frame-rate mismatch. Pick one per run and check fps with ffprobe on the first clip.
- video-filter dim amount 0 or 1.2 is refused: the (0, 1] range
Dim amount takes values above 0 up to 1. Zero, negatives and 1.2 return video_filter_amount_out_of_range. What it does, and how to lift a dark clip.
- video-filter invalid_filtergraph: every reason and the fix for each
A video-filter filtergraph is refused for eight reasons, from a [0:v] label to a quote character. What each one means and how to rewrite the graph so it passes.
Written by Sume