A speaker withdrew voice consent: finding every narration that used it
Revocation is a clean-up job. Build a manifest of voice, model and job id for each narration, then regenerate with another voice and join audio on Sume.

When a speaker withdraws consent for a synthetic voice, the work is finding every file that used it and replacing those files. You can only do that quickly if you recorded, for each narration, which voice and model made it and where the file went. Sume's job records carry the model for each job (job.model equals the routable id you requested, such as sonic-3.6), and its docs note that a thread can read the voice and model of each narration job. The voice-to-delivery link is yours to keep. A job list is not an index of voice usage.
This post covers the order of operations, a short script that turns a manifest into a replacement plan, and what a replacement costs. It is not legal advice; what you must delete and by when depends on the agreement and the place.
Order of operations
- Stop new use first: disable the voice in your own tooling so nobody generates with it while you work.
- Pull the manifest rows for that voice: job id, model, transcript id, destination and date delivered.
- Group by destination, because ads, product pages and a support line are replaced by different people.
- Regenerate each transcript with an approved voice and the same text, and rejoin audio where the video used several takes.
- Re-render or re-mux, replace the live asset, and mark the old file for deletion with its owner.
- Record what you did and when, so the answer to the speaker is a list, not a promise.
Know what the platform will and will not tell you
Under Jobs and results, an API key reads the jobs its own member created in the key's workspace, and a thread_id filter narrows a list without widening access. That is useful for scoping a clean-up to one project. It is not a promise that you can search every job by voice, so do not plan the clean-up around it. Plan around your own manifest.
The source-bound TTS contract helps one narrow part: a job made from an accepted script exposes a server-owned transcript_receipt with the job id, revision, sentence ids and the submitted transcript hash. That lets you regenerate exactly the same sentences, and it proves which text was spoken. It does not name the voice.
From manifest to plan
Counting characters per row tells you the cost before you start. At Sume's TTS Router rate of $47.50 per 1M characters, regenerating 40 narrations of 450 characters each is 18,000 characters, or about $0.86.
| Job id | Model | Voice | Delivered to | Action |
|---|---|---|---|---|
| job_a1 | sonic-3.6 | voice-17 | Paid social, week 40 | Regenerate, re-mux |
| job_a2 | sonic-3.6 | voice-17 | Product page video | Regenerate, re-render |
| job_b7 | sonic-3.5 | voice-22 | Support hold message | Keep, other voice |
The replacement plan, runnable
This reads a manifest and prints, for one voice, the rows to redo and the total characters and cost at Sume's rate.
manifest = [
{"job": "job_a1", "voice": "voice-17", "chars": 450, "dest": "paid social"},
{"job": "job_a2", "voice": "voice-17", "chars": 620, "dest": "product page"},
{"job": "job_b7", "voice": "voice-22", "chars": 300, "dest": "support line"},
]
revoked = "voice-17"
RATE = 47.50 # dollars per 1M characters, Sume TTS Router
todo = [m for m in manifest if m["voice"] == revoked]
by_dest = {}
for m in todo:
by_dest.setdefault(m["dest"], []).append(m["job"])
chars = sum(m["chars"] for m in todo)
print("jobs to redo:", len(todo))
print("by destination:", by_dest)
print("characters:", chars, " cost: $%.2f" % (chars * RATE / 1_000_000))
Prevent the next one
The clean-up is cheap if you prepared and slow if you did not. Three habits make the next one easy. Write the voice id into your own ledger at submit time, next to the job id, so the record exists before anyone asks. Put an expiry date on every voice and review the list monthly. And keep scripts as source-bound transcripts where you can: the accepted revision and sentence ids mean you can regenerate unchanged text exactly, and the verify step can reuse unchanged jobs instead of regenerating them.
If the speaker's agreement ends on a date you knew about, schedule the replacement before the date, not after. Regenerating forty short narrations is an afternoon, and a missed deadline is a conversation nobody wants.
Joining the replacement takes
If a video used several takes, you can rebuild the spine without re-synthesizing the good ones: timeline audio concatenates up to 20 Sume-hosted parts sample by sample at $0.01 flat per job and returns segments[] with offsets so you can re-base the video slots. Replace only the takes made with the withdrawn voice, keep the rest, and re-render.
Sources
Related posts
More in Use cases
- Splice an AI ending onto a Short and keep the original audio
Gemini Omni clips come with their own sound. To keep your Short's voice and music, detach both audio tracks and join them as audio parts in one Timeline render.
- Split a Short's audio at a cut point: timeline-audio ranges
YouTube Create has a split tool for audio. Sume's timeline-audio endpoint splits one file into up to 20 ranges, or joins up to 20 parts, for $0.01 a job.
- Sports replay Shorts: YouTube allows them if you explain the moves
YouTube lists sports replays where you explain what the competitor did as allowed. How to build a vertical breakdown Short from your own match footage.
- Spotify Clips ended: upload a 1080p full-length video instead
Spotify ended new Clips uploads June 17, 2026 and now takes full-length 16:9 video. See the spec and build a 1920x1080 file from your song with Sume Timeline.
Written by Sume