Fix one sentence in a 30-minute narration: 1 cent plus a re-render
Re-recording a 180-character sentence is 1 cent of Sume TTS. Swapping it into a 30-minute Timeline spine means a $3.00 re-render. Read model_id and voice first.

Re-recording one 180-character sentence costs 1 cent on Sume TTS, because 180 x 47.5 = 8,550 micro-dollars rounds up to a single cent. The expensive part is the re-render: a 30-minute Timeline render is 30 x $0.10 = $3.00. So the fix costs $3.01, and the way to keep it there is to read the original job's settings, so the retake matches on the first try.
Copy the settings from the job
A completed text-to-speech job records how Sume made the audio: model_id (the engine), voice (mode id and the id), language, output_format, and the synthesis settings generation_config and speed. Each setting is null when the request did not send it. Send the same values with the new sentence. If the original request sent no speed, send none.
If you used sonic-latest, the engine behind it may have moved since the first take. Pin the id you read from model_id, such as sonic-3.6, for the retake.
| Step | Math | Cost |
|---|---|---|
| TTS retake, 180 characters | 180 x 47.5 = 8,550 micro-dollars, round up | $0.01 |
| Timeline re-render, 30 minutes | 30 x $0.10 | $3.00 |
| Total | $3.01 |
Swap the slice, not the file
Timeline audio.parts[] accepts up to 20 gapless slices, each with a url, an optional source_in and an optional duration. Cut the original narration around the bad sentence with Timeline audio split (up to 20 ranges, each with start and optional end), then list the head slice, the new sentence, and the tail slice in order. The join is in the sample domain: no re-synthesis and no silence at the seams.
Check the seam by ear. A retake can have slightly different pacing, and a sentence boundary is the best place to cut.
Plan the render first
The unbilled plan call runs the schema and compiler checks and reports billable_minutes and an estimated cost without creating a job. Use it to confirm that the new total length still rounds to the same number of minutes. If the retake makes the spine 30 minutes 4 seconds, the render is rated as 31 minutes ($3.10), and the cap is 1,800 seconds, so it would fail.
When it is cheaper to redo the whole thing
Never. The TTS retake is a cent, and the render is the same as it was the first time. The cost of a fix is almost entirely the render, so batch fixes: if a client sends five corrections, retake all five sentences (5 cents at most) and re-render once ($3.00), not five times ($15.00).
Collect corrections before you render. Use the plan call to confirm the new total still fits in the same billed minute.
Sources
Related posts
More in Developers
- Four music variants in parallel: Python asyncio, one key per take
Submit four Music Router jobs at once from Python with one Idempotency-Key each, poll status concurrently, and pay 4 x $0.125 = $0.50. Code that runs.
- Free plan accepts 6 paid jobs: submit 7 and read queue_full
Free allows 1 processing job and 5 queued, so 6 are accepted. A Python script submits 7 in parallel and prints which are accepted and which get 429 queue_full.
- Free, Pro, Startup, Scale: processing seats, queue slots, full hold
Sume's concurrency by plan, queue capacity max(3, 5 x concurrency), accepted job capacity, and the balance reserved if every slot holds a 10 s clip.
- Gate a Format run on TikTok limits: artifact size, length, pixels
Check a Sume Format run's artifacts[] for size, duration and dimensions against TikTok's non-Spark limits before upload. A Python gate under 30 lines.
Written by Sume