AI avatar video editing: what you can change after rendering

AI avatar video editing: trims, captions, music, B-roll, and joins work on the finished MP4; new words, a new voice, or a new look need a re-render.

5 min readSume
All posts

AI avatar video editing depends on what you want to change. Cutting the start or end, cropping, adding captions, a music bed, or B-roll cutaways, and joining clips are edits on the finished MP4. New words, a new voice, a new avatar or setting, or a wider frame mean rendering again, because those are baked into the video. With Sume, edits run on the video's own media.sume.com URL through its media tools, and a re-render can start from a preview so you check it first.

Sume facts come from the Generate avatar video, Video trim, and Timeline 1.0 docs, read on 2026-09-28; anything called current behavior is read from Sume's code. For every editing endpoint in one map, see Video editing API: which endpoint.

Which changes are edits, and which need a new render?

Sort the change first. For the middle cut, Remove part of a video by API shows the render.

From Video trim, Video captions, Video filter, Timeline 1.0, and Generate avatar video, read 2026-09-28.
ChangeEdit or re-renderHow on Sume
Cut the start or the endEditPOST /v1/video-trim with start and end or duration
Remove a line in the middleEditA Timeline 1.0 render of the ranges you keep
Burn in captionsEditPOST /v1/video-captions on the finished URL
Add music or B-roll cutawaysEditTimeline 1.0: a soundtrack, and slots over the voice
Crop to a narrower frame or dim itEditPOST /v1/video-filter with a crop or dim op
Join several clipsEditTimeline 1.0
New words, a new avatar, or a new settingRe-renderA new talking video, checked on a preview first
A different voiceRe-renderTTS 1.0 speech, then VEED Fabric 1.0 lip sync of the avatar's still, without the original scene
A wider shape, such as 9:16 to 16:9Re-renderA new talking video with the new aspect_ratio

How do I edit the finished avatar video?

Use its own URL. Completed results are public media.sume.com artifacts, and the media tools read only your workspace's media.sume.com files, such as an earlier job's output. Each tool returns a new file and leaves the source untouched, so you can chain them:

  • Trim takes video_url, start, and exactly one of end or duration. A cut runs 0.2–900 seconds and keeps the audio by default.
  • Caption with POST /v1/video-captions. Today it refuses a source over 60 seconds or one with no audio stream, so caption each clip before you join clips into something longer.
  • Every edit is a job: poll GET /v1/jobs/{id}/status until the job is terminal, then, if it completed, read the new file's URL from GET /v1/jobs/{id}/result.
curl -X POST https://api.sume.com/v1/video-trim \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: welcome-v1-trim-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/welcome.mp4",
    "start": 0,
    "end": 12.5
  }'

Why does the avatar's voice disappear in a Timeline edit?

In current code, a Timeline render takes its sound only from the audio spine and an optional soundtrack; each clip's own audio is dropped. Before you add music or B-roll, detach the avatar's voice with POST /v1/audio-detach, which returns a wav by default, and use it as the spine. Slots can then cut between the avatar and B-roll while the words keep playing, as in Add B-roll to a talking-head avatar video.

Music goes in the same render as a Sume-hosted soundtrack bed; Talking head video background music covers levels, fades, and ducking.

How do I fix a word or change the script?

Render again. The avatar-video routes only list and read finished videos, and Sume's lip sync takes a still or an avatar, not a video, so the words can't be changed in place. To check the new version before paying for it:

  • Create an avatar video preview with the new script and approve its first-frame still, then call generate-video.
  • On a preview, quality can change at generate-video, but script, video_inputs, avatar_handle, scene, and aspect_ratio need a new preview.
  • For a long video joined from parts, re-render only the part with the mistake and run the join again with the new file in its place.
  • Use a new Idempotency-Key for the new request: the same key with a different body answers 409 idempotency_conflict.

What does editing cost?

Trim, audio detach, and caption jobs are priced per job; confirm them in GET /v1/catalog. A Timeline render is listed at $0.10 per output minute. A re-render bills per second of video again: $0.184/s standard, $0.245/s plus, $0.55/s max (no product image). All are plus a 5.5% agent fee by default.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume