How do I fix one sentence in finished AI narration without redoing it?
Retake just the wrong sentence with a one-cent TTS job, then splice it into the original file with a $0.01 Timeline audio concat using source_in and duration.

To fix one sentence in finished AI narration, ask for sentence timings when you first generate the file, retake only the wrong sentence as a new TTS job, and splice it in with a Timeline audio concat that uses the original file twice, once before the sentence and once after. A 150-character retake costs $0.01 and the concat is a flat $0.01, so the fix is about 2 cents instead of paying for the whole narration again.
The one condition is that the splice works only if the new line sounds like it came from the same take. Use the same voice, language, speed and emotion text as the original, and treat any change of those as a full re-record.
Ask for timings at the start
The splice needs to know where each sentence begins and ends. On the first generation set timestamps.words and segmentation.mode to sentence; segmentation requires word timestamps. The result carries the segments with start times, and with wav or raw output you also get a per-segment audio URL. Keep that result next to the audio, because the offsets are what you edit against.
Without timestamps you can still find the cut by ear, but you lose the ability to script the fix. Adding the two fields later means generating again.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: narration-ch2-v1" \
-d @- <<JSON
{
"transcript": "$(cat chapter2.txt)",
"voice": { "id": "$VOICE_ID" },
"language": "en",
"timestamps": { "words": true },
"segmentation": { "mode": "sentence" },
"output_format": { "container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100 }
}
JSONRetake the sentence
Generate the corrected sentence as its own job with the same settings. Use a new idempotency key, for example narration-ch2-s14-v2, so it is billed once and a retry is safe. Billing rounds each job up to a whole cent with a one-cent minimum, so a short fix costs one or two cents however short it is.
Listen to the line alone first. If its pace or tone is off, adjust generation_config.speed or the emotion text and retake once more; each retake is another cent or two, not a re-render of the chapter.
Splice with source_in and duration
Timeline audio concat parts each take a url, an optional source_in (the start offset into that file) and an optional duration. To replace sentence 14, which runs from 61.2 s to 66.8 s in the original, send three parts: the original from 0 for 61.2 seconds, the retake, and the original again from 66.8 seconds. The join is sample-domain with no gap and no re-synthesis, so the untouched speech is bit-for-bit what you had.
curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: narration-ch2-spliced-v2" \
-d '{
"operation": "concat",
"parts": [
{ "url": "'"$ORIGINAL"'", "source_in": 0, "duration": 61.2 },
{ "url": "'"$RETAKE"'" },
{ "url": "'"$ORIGINAL"'", "source_in": 66.8 }
]
}'When a splice is the wrong tool
A splice repairs a sentence, not a performance. If the wrong sentence sits in the middle of a rising passage, or the fix changes the length a lot, the join can sound abrupt. In that case retake the paragraph around it: the neighbouring sentences are in the segment list, so the splice points are just further apart.
If the error is in many places, for example a product name that changed throughout, fix the script and regenerate the whole file with a pronunciation dictionary attached. Patching twenty sentences costs twenty jobs, plus staged concats, and the result is easier to get wrong than one clean generation.
Finally, treat the spliced file as a new version. Name it, store the cut points next to it, and keep the original for the next fix.
Cost and limits
All parts must be this workspace's media.sume.com audio with the same channel layout, so use the original TTS result URL rather than a re-encoded copy. If you plan to splice repeatedly, keep wav: mp3 adds encoder padding that can show up as a tiny gap at the seams. Concat takes 1 to 20 parts, so a chapter with many fixes is spliced in stages.
- Keep the original file and its segment list; never overwrite them.
- Retake with identical settings, or the seam will be audible.
- Check the splice by listening to five seconds either side of the join.
| Approach | What you pay for | Cost |
|---|---|---|
| Regenerate the whole narration | 15,000 characters at $0.0475 per 1,000 | $0.72 |
| Retake and splice | 150 characters plus one concat | $0.02 |
| Cartesia list overage, Scale plan | $38 per 1M credits at 1 credit per character, 15,000 credits | $0.570 |
Sources
Related posts
More in Developers
- Get transcript text from a captioned video: caption jobs return none
A Sume caption job returns the burned video, not the transcript. For text and word times run STT or video inspect with transcribe. Prices and a recipe.
- Handle every Sume API error with one switch on next_action
Sume errors share one envelope. Branch on next_action, retryable and retry_after_seconds, and your client handles new codes without a code change. JS sample.
- Hindi speech to text API: Sume STT with language_code hi
Transcribe Hindi audio with Sume STT: send language_code hi, check the reported language, and review code-mixed speech. $0.01 per audio minute.
- A 3-minute Timeline render: poll job status, don't sleep a fixed time
Timeline renders are async by default. Submit with an Idempotency-Key, poll /v1/jobs/:id/status until terminal, then read /result. Cost: 3 minutes is $0.30.
Written by Sume