How do I fix one sentence in finished AI narration without redoing it?

Retake just the wrong sentence with a one-cent TTS job, then splice it into the original file with a $0.01 Timeline audio concat using source_in and duration.

4 min readSume
All posts

To fix one sentence in finished AI narration, ask for sentence timings when you first generate the file, retake only the wrong sentence as a new TTS job, and splice it in with a Timeline audio concat that uses the original file twice, once before the sentence and once after. A 150-character retake costs $0.01 and the concat is a flat $0.01, so the fix is about 2 cents instead of paying for the whole narration again.

The one condition is that the splice works only if the new line sounds like it came from the same take. Use the same voice, language, speed and emotion text as the original, and treat any change of those as a full re-record.

Ask for timings at the start

The splice needs to know where each sentence begins and ends. On the first generation set timestamps.words and segmentation.mode to sentence; segmentation requires word timestamps. The result carries the segments with start times, and with wav or raw output you also get a per-segment audio URL. Keep that result next to the audio, because the offsets are what you edit against.

Without timestamps you can still find the cut by ear, but you lose the ability to script the fix. Adding the two fields later means generating again.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: narration-ch2-v1" \
  -d @- <<JSON
{
  "transcript": "$(cat chapter2.txt)",
  "voice": { "id": "$VOICE_ID" },
  "language": "en",
  "timestamps": { "words": true },
  "segmentation": { "mode": "sentence" },
  "output_format": { "container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100 }
}
JSON

Retake the sentence

Generate the corrected sentence as its own job with the same settings. Use a new idempotency key, for example narration-ch2-s14-v2, so it is billed once and a retry is safe. Billing rounds each job up to a whole cent with a one-cent minimum, so a short fix costs one or two cents however short it is.

Listen to the line alone first. If its pace or tone is off, adjust generation_config.speed or the emotion text and retake once more; each retake is another cent or two, not a re-render of the chapter.

Splice with source_in and duration

Timeline audio concat parts each take a url, an optional source_in (the start offset into that file) and an optional duration. To replace sentence 14, which runs from 61.2 s to 66.8 s in the original, send three parts: the original from 0 for 61.2 seconds, the retake, and the original again from 66.8 seconds. The join is sample-domain with no gap and no re-synthesis, so the untouched speech is bit-for-bit what you had.

curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: narration-ch2-spliced-v2" \
  -d '{
    "operation": "concat",
    "parts": [
      { "url": "'"$ORIGINAL"'", "source_in": 0, "duration": 61.2 },
      { "url": "'"$RETAKE"'" },
      { "url": "'"$ORIGINAL"'", "source_in": 66.8 }
    ]
  }'

When a splice is the wrong tool

A splice repairs a sentence, not a performance. If the wrong sentence sits in the middle of a rising passage, or the fix changes the length a lot, the join can sound abrupt. In that case retake the paragraph around it: the neighbouring sentences are in the segment list, so the splice points are just further apart.

If the error is in many places, for example a product name that changed throughout, fix the script and regenerate the whole file with a pronunciation dictionary attached. Patching twenty sentences costs twenty jobs, plus staged concats, and the result is easier to get wrong than one clean generation.

Finally, treat the spliced file as a new version. Name it, store the cut points next to it, and keep the original for the next fix.

Cost and limits

All parts must be this workspace's media.sume.com audio with the same channel layout, so use the original TTS result URL rather than a re-encoded copy. If you plan to splice repeatedly, keep wav: mp3 adds encoder padding that can show up as a tiny gap at the seams. Concat takes 1 to 20 parts, so a chapter with many fixes is spliced in stages.

  • Keep the original file and its segment list; never overwrite them.
  • Retake with identical settings, or the seam will be audible.
  • Check the splice by listening to five seconds either side of the join.
Fixing one 150-character sentence in a 15,000-character narration; Sume TTS 1.0 and Timeline audio catalog rates as of 2026-10-07, Cartesia list overage rate read 2026-10-07.
ApproachWhat you pay forCost
Regenerate the whole narration15,000 characters at $0.0475 per 1,000$0.72
Retake and splice150 characters plus one concat$0.02
Cartesia list overage, Scale plan$38 per 1M credits at 1 credit per character, 15,000 credits$0.570

Sources

Related posts

More in Developers

All Developers posts

Written by Sume