A speaker withdrew voice consent: finding every narration that used it

Revocation is a clean-up job. Build a manifest of voice, model and job id for each narration, then regenerate with another voice and join audio on Sume.

6 min readSume
All posts

When a speaker withdraws consent for a synthetic voice, the work is finding every file that used it and replacing those files. You can only do that quickly if you recorded, for each narration, which voice and model made it and where the file went. Sume's job records carry the model for each job (job.model equals the routable id you requested, such as sonic-3.6), and its docs note that a thread can read the voice and model of each narration job. The voice-to-delivery link is yours to keep. A job list is not an index of voice usage.

This post covers the order of operations, a short script that turns a manifest into a replacement plan, and what a replacement costs. It is not legal advice; what you must delete and by when depends on the agreement and the place.

Order of operations

  • Stop new use first: disable the voice in your own tooling so nobody generates with it while you work.
  • Pull the manifest rows for that voice: job id, model, transcript id, destination and date delivered.
  • Group by destination, because ads, product pages and a support line are replaced by different people.
  • Regenerate each transcript with an approved voice and the same text, and rejoin audio where the video used several takes.
  • Re-render or re-mux, replace the live asset, and mark the old file for deletion with its owner.
  • Record what you did and when, so the answer to the speaker is a list, not a promise.

Know what the platform will and will not tell you

Under Jobs and results, an API key reads the jobs its own member created in the key's workspace, and a thread_id filter narrows a list without widening access. That is useful for scoping a clean-up to one project. It is not a promise that you can search every job by voice, so do not plan the clean-up around it. Plan around your own manifest.

The source-bound TTS contract helps one narrow part: a job made from an accepted script exposes a server-owned transcript_receipt with the job id, revision, sentence ids and the submitted transcript hash. That lets you regenerate exactly the same sentences, and it proves which text was spoken. It does not name the voice.

From manifest to plan

Counting characters per row tells you the cost before you start. At Sume's TTS Router rate of $47.50 per 1M characters, regenerating 40 narrations of 450 characters each is 18,000 characters, or about $0.86.

Example manifest rows and the action each needs; fields are a suggested minimum, read 2026-10-03 against Sume's job model fields.
Job idModelVoiceDelivered toAction
job_a1sonic-3.6voice-17Paid social, week 40Regenerate, re-mux
job_a2sonic-3.6voice-17Product page videoRegenerate, re-render
job_b7sonic-3.5voice-22Support hold messageKeep, other voice

The replacement plan, runnable

This reads a manifest and prints, for one voice, the rows to redo and the total characters and cost at Sume's rate.

manifest = [
    {"job": "job_a1", "voice": "voice-17", "chars": 450, "dest": "paid social"},
    {"job": "job_a2", "voice": "voice-17", "chars": 620, "dest": "product page"},
    {"job": "job_b7", "voice": "voice-22", "chars": 300, "dest": "support line"},
]
revoked = "voice-17"
RATE = 47.50  # dollars per 1M characters, Sume TTS Router

todo = [m for m in manifest if m["voice"] == revoked]
by_dest = {}
for m in todo:
    by_dest.setdefault(m["dest"], []).append(m["job"])

chars = sum(m["chars"] for m in todo)
print("jobs to redo:", len(todo))
print("by destination:", by_dest)
print("characters:", chars, " cost: $%.2f" % (chars * RATE / 1_000_000))

Prevent the next one

The clean-up is cheap if you prepared and slow if you did not. Three habits make the next one easy. Write the voice id into your own ledger at submit time, next to the job id, so the record exists before anyone asks. Put an expiry date on every voice and review the list monthly. And keep scripts as source-bound transcripts where you can: the accepted revision and sentence ids mean you can regenerate unchanged text exactly, and the verify step can reuse unchanged jobs instead of regenerating them.

If the speaker's agreement ends on a date you knew about, schedule the replacement before the date, not after. Regenerating forty short narrations is an afternoon, and a missed deadline is a conversation nobody wants.

Joining the replacement takes

If a video used several takes, you can rebuild the spine without re-synthesizing the good ones: timeline audio concatenates up to 20 Sume-hosted parts sample by sample at $0.01 flat per job and returns segments[] with offsets so you can re-base the video slots. Replace only the takes made with the withdrawn voice, keep the rest, and re-render.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume