Fix one sentence in an AI avatar video without a full re-render

ElevenLabs can regenerate only edited dubbing regions. A Sume avatar video is one job, so the fix is to keep clips short and join them: the cost math.

5 min readSume
All posts

You cannot patch one sentence inside a finished Sume avatar video; the render is one job, and changing the script means a new job for the whole clip. The way to make that cheap is to render in short pieces, so a typo costs you one piece rather than a minute. ElevenLabs' Dubbing v2 API works the other way: its changelog entry dated 2026-08-10 says you can edit individual transcript segments or translations and regenerate only the regions that changed.

That is a real difference, and it decides how you structure a talking-head project.

What a re-render costs

At Sume's standard rate of $0.184 per second with no product image, the cost of fixing one line depends on how big the unit of work was.

Arithmetic from the Sume rate card (standard, no product image, $0.184 per second); confirm with GET /v1/catalog. ElevenLabs behavior from its changelog, read 2026-10-01.
How you renderedUnit to redo for one wrong sentenceCost on standard
One 60 s videoThe whole 60 s$11.04
Three 20 s clipsOne 20 s clip$3.68
Six 10 s clipsOne 10 s clip$1.84
Dubbing v2 project (ElevenLabs)Only the edited regions, per its changelogNot stated in the changelog

A workflow that stays cheap to fix

Keep the audio side in mind too. Timeline audio can join Sume-hosted audio parts into one gapless file for $0.01 a job, but that is for audio you already have, not for patching a rendered face.

  • Write the script as scenes of one idea each and render each as its own talking-video job with the same avatar_handle and the same scene so the look matches.
  • Use an Idempotency-Key that names the scene and its script version, such as ep12-scene3-v2. A retry of the same version returns the same job; a rewrite gets a new key and a new job.
  • Approve framing before you pay for motion with an avatar video preview. Preview stills are reused at generate-video, and you can change only the final quality tier there.
  • Join the approved pieces on the Timeline 1.0 surface, which sequences clips; the docs list it as the surface for sequencing several clips. A render is $0.10 per output minute, rounded up.

Submitting a single scene

Each piece is an ordinary request. The total of a scene must still estimate to at least 4 seconds.

curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: ep12-scene3-v2" \
  -d '{
    "avatar_handle": "studio_presenter",
    "script": "Refunds post within five business days, not ten.",
    "quality": "standard"
  }'

Limits

Cuts between separately rendered clips can show small differences in pose or lighting, because each clip is generated on its own. Current execution supports one resolved avatar per final video and expects scene backgrounds to resolve to one shared scene. Short clips under 4 seconds are not accepted. The ElevenLabs behavior concerns dubbing audio, not avatar video, so it is a comparison of workflow, not of output.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume