Shorts conversational editing vs API rough cuts with Trim
YouTube added conversational AI editing to Shorts. For repeatable rough cuts, Sume's video trim cuts a range from one clip into a new MP4 for $0.02.

Conversational editing in the YouTube app is a good fit for one-off edits by hand. For rough cuts you repeat across many clips, Sume's video trim takes one clip and a range, returns a new MP4 holding only that range, and costs $0.02 per job. Each cut is a recorded API call with an id.
What YouTube announced
Reports on Made on YouTube 2026 from Sep 23, 2026 list conversational AI editing in Shorts, expanded likeness detection, real-time live auto-dubbing and Shorts series on TV. They do not give limits or availability, so check YouTube for specifics.
What Trim does
POST /v1/video-trim takes one media.sume.com clip, a start in seconds, and exactly one of end or duration. The source is untouched and a new artifact is returned. The video must already be in your workspace; there is no open-internet fetch, so import first with POST /v1/media-imports. Idempotency-Key is required.
curl -X POST https://api.sume.com/v1/video-trim \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: video-trim-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"start": 2,
"duration": 8
}'Options and limits
Defaults are precision: "exact", a frame-accurate re-encode, and audio: "keep". precision: "keyframe" is a stream copy whose cut may start a GOP early; re-base against actual_start_seconds. An output conform of width, height and fps works with exact only. The source may be up to 1800 seconds and the output 0.2 to 900 seconds. An end past the source clamps, and the result warns trim_clamped_to_source.
| Question | Shorts conversational editing | Sume video trim |
|---|---|---|
| How you drive it | Conversation in the app | One API request |
| Repeatable across many clips | Not described in the announcement | Yes, same request with new values |
| Record of each edit | Not described | Job id and events |
| Cost | Not stated in the summary | $0.02 per job; confirm in GET /v1/catalog |
Where an avatar clip fits
If the clip is a talking presenter, make it with POST /v1/avatar-1.0/talking-video (scripts of 4 to 60 seconds), then trim or assemble it. Timeline 1.0 joins clips over one audio spine; captions can be burned last. Poll GET /v1/jobs/{id}/status for each job and read the new MP4 from the result.
Sources
Related posts
More in Comparisons
- Speech Arena Elo 1,319: what the score means
Eleven v4 reportedly leads Artificial Analysis' Provider Voice Arena at Elo 1,319. What an arena Elo tells you and how to run your own blind test.
- Spotify AI covers and remixes: licensed tool vs an original score
Spotify and UMG announced a paid add-on for fan covers and remixes. A licensed remix tool is not an original video score. Here is the Sume route for the second.
- Stable Diffusion alternatives in 2026: open weights or hosted API
SD 3.5 is still Stability's newest flagship image model. FLUX.2 [klein], Qwen-Image 2.0 and Z-Image Turbo are the open options. How to pick between them.
- Suno Speech beta: voice and music in one pass, or separate tracks?
Suno's Speech beta makes voice and music in one track. Its blog lists wandering accents and long pauses. When to prefer separate TTS, music and a timeline mix.
Written by Sume