AI avatar video editing: what you can change after rendering
AI avatar video editing: trims, captions, music, B-roll, and joins work on the finished MP4; new words, a new voice, or a new look need a re-render.
AI avatar video editing depends on what you want to change. Cutting the start or end, cropping, adding captions, a music bed, or B-roll cutaways, and joining clips are edits on the finished MP4. New words, a new voice, a new avatar or setting, or a wider frame mean rendering again, because those are baked into the video. With Sume, edits run on the video's own media.sume.com URL through its media tools, and a re-render can start from a preview so you check it first.
Sume facts come from the Generate avatar video, Video trim, and Timeline 1.0 docs, read on 2026-09-28; anything called current behavior is read from Sume's code. For every editing endpoint in one map, see Video editing API: which endpoint.
Which changes are edits, and which need a new render?
Sort the change first. For the middle cut, Remove part of a video by API shows the render.
| Change | Edit or re-render | How on Sume |
|---|---|---|
| Cut the start or the end | Edit | POST /v1/video-trim with start and end or duration |
| Remove a line in the middle | Edit | A Timeline 1.0 render of the ranges you keep |
| Burn in captions | Edit | POST /v1/video-captions on the finished URL |
| Add music or B-roll cutaways | Edit | Timeline 1.0: a soundtrack, and slots over the voice |
| Crop to a narrower frame or dim it | Edit | POST /v1/video-filter with a crop or dim op |
| Join several clips | Edit | Timeline 1.0 |
| New words, a new avatar, or a new setting | Re-render | A new talking video, checked on a preview first |
| A different voice | Re-render | TTS 1.0 speech, then VEED Fabric 1.0 lip sync of the avatar's still, without the original scene |
| A wider shape, such as 9:16 to 16:9 | Re-render | A new talking video with the new aspect_ratio |
How do I edit the finished avatar video?
Use its own URL. Completed results are public media.sume.com artifacts, and the media tools read only your workspace's media.sume.com files, such as an earlier job's output. Each tool returns a new file and leaves the source untouched, so you can chain them:
- Trim takes
video_url,start, and exactly one ofendorduration. A cut runs 0.2–900 seconds and keeps the audio by default. - Caption with
POST /v1/video-captions. Today it refuses a source over 60 seconds or one with no audio stream, so caption each clip before you join clips into something longer. - Every edit is a job: poll
GET /v1/jobs/{id}/statusuntil the job is terminal, then, if it completed, read the new file's URL fromGET /v1/jobs/{id}/result.
curl -X POST https://api.sume.com/v1/video-trim \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: welcome-v1-trim-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/welcome.mp4",
"start": 0,
"end": 12.5
}'Why does the avatar's voice disappear in a Timeline edit?
In current code, a Timeline render takes its sound only from the audio spine and an optional soundtrack; each clip's own audio is dropped. Before you add music or B-roll, detach the avatar's voice with POST /v1/audio-detach, which returns a wav by default, and use it as the spine. Slots can then cut between the avatar and B-roll while the words keep playing, as in Add B-roll to a talking-head avatar video.
Music goes in the same render as a Sume-hosted soundtrack bed; Talking head video background music covers levels, fades, and ducking.
How do I fix a word or change the script?
Render again. The avatar-video routes only list and read finished videos, and Sume's lip sync takes a still or an avatar, not a video, so the words can't be changed in place. To check the new version before paying for it:
- Create an avatar video preview with the new script and approve its first-frame still, then call
generate-video. - On a preview,
qualitycan change atgenerate-video, butscript,video_inputs,avatar_handle,scene, andaspect_rationeed a new preview. - For a long video joined from parts, re-render only the part with the mistake and run the join again with the new file in its place.
- Use a new
Idempotency-Keyfor the new request: the same key with a different body answers409 idempotency_conflict.
What does editing cost?
Trim, audio detach, and caption jobs are priced per job; confirm them in GET /v1/catalog. A Timeline render is listed at $0.10 per output minute. A re-render bills per second of video again: $0.184/s standard, $0.245/s plus, $0.55/s max (no product image). All are plus a 5.5% agent fee by default.
Sources
Related posts
More in Sume Avatar 1.0
- HeyGen Avatar 4 vs 5: Avatar IV and Avatar V compared
HeyGen Avatar IV is the default engine for every avatar type; Avatar V is opt-in for eligible Digital Twins only. Parameters, eligibility and cost.
- HeyGen video translation API: modes, languages, cost
HeyGen's video translation API is POST /v3/video-translations: a video URL plus target languages, in Speed or Precision mode, billed per minute.
- HeyGen vs Argil: avatar APIs, avatar creation and pricing
HeyGen and Argil both turn a script into an avatar video by API. HeyGen bills pay-as-you-go dollars per minute; Argil bills plan credits per minute.
- HeyGen vs Creatify: avatar API, ad tools and pricing
HeyGen's API centers on avatars, voice and translation, paid with pay-as-you-go credits. Creatify's centers on video ads, sold as monthly API plans.
Written by Sume