HeyGen scene-scoped edits vs fixing one scene in a Sume avatar
HeyGen's Video Agent can edit one scene and keep the rest. Sume has no scene-edit call: structural changes need a new preview; only quality changes late.

HeyGen's Video Agent now accepts a scene-scoped edit plan that leaves untouched scenes alone. Sume's avatar video API has no equivalent edit call: once a video is rendered, changing the script means a new render, and the cheap place to catch mistakes is the preview stage before it.
This post lays out what HeyGen documents, what Sume documents, and the workflow that follows for a script that keeps changing.
What did HeyGen add for editing one scene?
The September 2026 changelog entry says POST /v3/video-agents/{session_id} accepts an edit_plan: one natural-language change per scene, applied together in a single turn. Untouched scenes keep their script and visuals. You need scene IDs from GET /v3/videos/{video_id}/scenes, and a request can carry up to 50 items, with validation that the scenes are consistent.
That is an edit against an existing video session.
What can you change on Sume without starting over?
Sume splits the work in two. An avatar video preview generates only the first-frame stills, one per input scene when you use video_inputs. You approve the composition, then call generate-video on the preview id and Sume reuses the stills.
Two things are explicitly cheap to change after approval. quality can be overridden at generate-video (standard, plus or max), because preview stills are tier-independent. And POST /v1/avatar-video-previews/{id}/regenerate reuses the stored request and refreshes only the first-frame stills.
What forces a new preview?
The docs list the structural fields: script, video_inputs, avatar_handle, scene and aspect_ratio. Change any of them and you create a new preview. There is no documented call that patches one scene of a finished avatar video and keeps the others.
| Change | Same preview? | How |
|---|---|---|
| Final render tier | Yes | quality on generate-video |
| Refresh first-frame stills | Yes | regenerate on the preview id |
| Script or video_inputs | No | Create a new preview |
| Avatar, scene, aspect ratio | No | Create a new preview |
How do you work when only one sentence changes?
Plan for it before you render. Keep each spoken beat in its own short job when the beats are likely to change, and join the finished clips later with Timeline 1.0, which handles sequencing and transitions. A fix then re-renders one short clip instead of the whole video.
The trade-off is real. Separate avatar videos are separate jobs, each with its own 4-60 second window and its own cost, and each is rendered on its own, so the framing of a shared scene is not guaranteed to match between them. Within a single multi-scene video, Sume resolves one avatar and one shared scene.
Is the preview step worth it for edits?
Yes when the risk is composition: framing, product placement, background. No when the risk is wording, because stills do not show what the avatar says. Review wording in your own tooling first, then preview, then render.
Captions are stored on preview create and burned in only at generate-video; preview stills are never captioned. For the related question of what happens after approval, read change the script after approving a preview, and for the whole-render angle fix one sentence in an AI avatar video.
Sources
Related posts
More in Comparisons
- HeyGen's 22 Spanish and 17 Arabic variants vs Sume's language field
HeyGen lists regional variants such as 22 Spanish and 17 Arabic. Sume's TTS takes one BCP-47 language string and does not publish a per-region variant list.
- Higgsfield AI video translator: 18 languages, lip sync, vs Sume
Higgsfield's video translator dubs into 18 languages and re-syncs lips. Sume has no one-call translator; here is what each does, and the Sume steps instead.
- Higgsfield UGC makes 4 generations per run: Sume queue by plan
Higgsfield says a UGC run can produce up to 4 generations. In Sume you submit one job per clip and plan limits set how many run and queue at once.
- Higgsfield UGC is capped at 15 seconds: longer clips on Sume
Higgsfield's UGC guide lists up to 15 seconds per video. Sume Avatar 1.0 accepts scripts of 4 to 60 seconds, and one avatar per final video. Limits and pricing.
Written by Sume