Paper edit by API: build a rough cut from transcript lines
Premiere 26.5 added Paper Edit. To do the same by API, transcribe with sentence segments, pick the lines you keep, then render those ranges as a Timeline.

To build a rough cut from transcript lines by API, run POST /v1/video-inspect with transcribe: true and sentence segmentation, choose the lines you keep, then send those ranges as video[] slots to POST /v1/timeline-1.0/render. Sume does not pick the lines for you; your code or an agent does.
Adobe's Premiere 26.5 announcement lists Paper Edit as a way to "assemble rough cuts directly from your transcript". This post covers the same idea through Sume's public routes, from the Video inspect and Timeline 1.0 docs, read 2026-09-30.
How do I get transcript lines with timing?
Send transcribe: true and segmentation: { "mode": "sentence" } to POST /v1/video-inspect. The result carries text, words[] and gapless sentence segments[] shaped like caption lines. Transcription is the only billed half: $0.01 per audio minute. Add frames: false if you do not want stills.
curl -X POST https://api.sume.com/v1/video-inspect \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: paper-edit-inspect-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/interview.mp4",
"frames": false,
"transcribe": true,
"segmentation": { "mode": "sentence" }
}'How do the chosen lines become a sequence?
Each kept line becomes one video[] slot: source_url is the interview, source_in is the line's in-point, duration is its length, and start is where it lands on the output. video[0].start must be 0 and later starts must increase. A Timeline also needs an audio spine with audio.duration_seconds (1 to 1800); to keep the interview's own sound, extract it with POST /v1/audio-detach and list the same ranges in audio.parts[] (each part takes url, source_in and duration; at most 20 parts).
What does Sume not do here?
The docs describe routes, not an editing UI, and none of the routes above chooses lines for you. Which lines to keep is a judgment that stays in your code or your agent.
| Step | Route | Cost in the docs |
|---|---|---|
| Transcribe with sentence lines | POST /v1/video-inspect | $0.01 per audio minute |
| Cut one kept range to its own MP4 | POST /v1/video-trim | $0.02 per job |
| Assemble slots with source_in | POST /v1/timeline-1.0/render | $0.10 per ceil(output minute) |
| Preview cost and length | POST /v1/timeline-1.0/plan | Unbilled |
Should I trim each line first or use source_in?
source_in on a slot reads a range straight out of the file, so a separate trim is not needed for a plain cut. Use video-trim only when you want each range as its own reusable MP4. Run /v1/timeline-1.0/plan before you render to see duration_seconds, segment_count and estimated_cost_usd_micros without creating a job.
Sources
Related posts
More in Developers
- Pipedream 30-second timeout: call Sume in async mode, not sync
Pipedream HTTP workflows time out at 30 seconds by default. Sume sync waits up to 30 seconds too, so submit async and take the result by webhook or poll.
- Pipedream 512KB body limit and Sume run webhook payloads
Pipedream limits HTTP trigger bodies to 512KB by default. Sume run webhooks can carry up to 1 MiB, so plan for a slim relay or the result_url fetch.
- Remove filler words from a talking video with an API
Sume has no one-call filler remover. Transcribe with video inspect for word timings, then cut the clean ranges with video trim at $0.02 per job.
- Remove filler words from a video by API: cut at word timestamps
Descript's API lists Remove Filler Words as an Underlord edit. the Sume docs list no such op; here is how to cut um and uh yourself from words[] and a Timeline.
Written by Sume