Fix one sentence in an AI avatar video without a full re-render
ElevenLabs can regenerate only edited dubbing regions. A Sume avatar video is one job, so the fix is to keep clips short and join them: the cost math.
You cannot patch one sentence inside a finished Sume avatar video; the render is one job, and changing the script means a new job for the whole clip. The way to make that cheap is to render in short pieces, so a typo costs you one piece rather than a minute. ElevenLabs' Dubbing v2 API works the other way: its changelog entry dated 2026-08-10 says you can edit individual transcript segments or translations and regenerate only the regions that changed.
That is a real difference, and it decides how you structure a talking-head project.
What a re-render costs
At Sume's standard rate of $0.184 per second with no product image, the cost of fixing one line depends on how big the unit of work was.
| How you rendered | Unit to redo for one wrong sentence | Cost on standard |
|---|---|---|
| One 60 s video | The whole 60 s | $11.04 |
| Three 20 s clips | One 20 s clip | $3.68 |
| Six 10 s clips | One 10 s clip | $1.84 |
| Dubbing v2 project (ElevenLabs) | Only the edited regions, per its changelog | Not stated in the changelog |
A workflow that stays cheap to fix
Keep the audio side in mind too. Timeline audio can join Sume-hosted audio parts into one gapless file for $0.01 a job, but that is for audio you already have, not for patching a rendered face.
- Write the script as scenes of one idea each and render each as its own
talking-videojob with the sameavatar_handleand the samesceneso the look matches. - Use an
Idempotency-Keythat names the scene and its script version, such asep12-scene3-v2. A retry of the same version returns the same job; a rewrite gets a new key and a new job. - Approve framing before you pay for motion with an avatar video preview. Preview stills are reused at
generate-video, and you can change only the final quality tier there. - Join the approved pieces on the Timeline 1.0 surface, which sequences clips; the docs list it as the surface for sequencing several clips. A render is $0.10 per output minute, rounded up.
Submitting a single scene
Each piece is an ordinary request. The total of a scene must still estimate to at least 4 seconds.
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: ep12-scene3-v2" \
-d '{
"avatar_handle": "studio_presenter",
"script": "Refunds post within five business days, not ten.",
"quality": "standard"
}'Limits
Cuts between separately rendered clips can show small differences in pose or lighting, because each clip is generated on its own. Current execution supports one resolved avatar per final video and expects scene backgrounds to resolve to one shared scene. Short clips under 4 seconds are not accepted. The ElevenLabs behavior concerns dubbing audio, not avatar video, so it is a comparison of workflow, not of output.
Sources
Related posts
More in Developers
- FLUX API 402, 403 and 503 errors, and Sume's equivalents
BFL returns 402 for credits, 403 for key permission, 503 for load. Sume returns 402 insufficient_credits and 503 provider_capacity_exceeded. What to do.
- Gemini API video upload limits vs how Sume takes media inputs
Gemini accepts video inline under 100 MB or through the File API up to 20 GB paid and 2 GB free. How Sume's video routes take URLs, and where they refuse.
- Gemini prefixItems tuple schema: Sume rejects it, use an object
Gemini lists prefixItems for tuple-like arrays. Sume's output_schema allowlist omits it and returns unsupported_keyword; model each slot as a named property.
- Gemini recursive schema with $ref "#": what Sume accepts instead
Gemini's docs show an org-chart schema that recurses with $ref "#". Sume rejects that root reference; recurse through a named $defs entry instead.
Written by Sume