Prove a voiceover read the approved script: the transcript receipt
A source-bound Sume TTS job carries a transcript receipt with the revision, sentence IDs and a SHA-256. It proves the text sent, not how it was pronounced.

If a reviewer asks whether the voiceover says what legal approved, a Sume TTS job built from a source-bound script can answer for the text. The job exposes a server-owned transcript_receipt with the job id, script revision, sentence ids, canonicalization version, the SHA-256 of the submitted transcript and input_integrity: "source_bound". A read-only verify step then checks selected jobs against the accepted script. What the receipt cannot prove is how the words sound.
With new speech models arriving at $15 to $22 per 1M characters from Microsoft (read 2026-10-03), more teams will generate narration at volume. Volume makes an audit trail worth having.
What a receipt covers
Raw transcript text, paths, client hashes and receipt fields are rejected inside transcript_source. You name sentences by id and the server resolves them, which is why a model cannot quietly edit a price line on the way to the voice.
| Question | Answer |
|---|---|
| Text input rule | Exactly one of transcript or transcript_source |
transcript_source fields | script_revision_id and sentence_ids only |
| Discover the accepted script | GET /v1/tts-1.0/source (tool tts_source_get, free) |
| Check jobs against it | POST /v1/tts-1.0/source/verify-spine (tool tts_source_verify_spine, free) |
| Receipt holds | Job id, revision, sentence ids, canonicalization version, transcript SHA-256, input_integrity |
| Receipt does not prove | Pronunciation, pacing, acoustic quality |
| TTS Router | Same contract, plus the required model |
A checker for your own records
Even without calling the verify route, you can keep your own ledger. Hash each approved sentence at approval time, hash what each job reports, and compare. The snippet below shows the shape of the comparison with local data, so you can adapt it to whatever your jobs return.
import hashlib
def sha(text):
return hashlib.sha256(text.encode("utf-8")).hexdigest()
approved = {
"s1": "Order by October 31.",
"s2": "Free shipping on orders over 50 dollars.",
}
# what you recorded for each generated job: sentence ids and the transcript hash
jobs = [
{"job": "job_a", "sentence_ids": ["s1"], "transcript_sha256": sha("Order by October 31.")},
{"job": "job_b", "sentence_ids": ["s2"], "transcript_sha256": sha("Free shipping on orders over 40 dollars.")},
]
for j in jobs:
expected = sha(" ".join(approved[i] for i in j["sentence_ids"]))
ok = expected == j["transcript_sha256"]
print(j["job"], "MATCH" if ok else "MISMATCH")
What to do with a mismatch, and with a missing receipt
Per the notes, mismatched text cannot be adopted, and a missing receipt on its own must never trigger automatic regeneration. For older jobs with no receipt, you name the sentence ids and the server compares the stored submitted text exactly after canonicalization; an exact match can be adopted with no writes, provider calls or regeneration. That is the cheap path, so try it before you pay to regenerate.
Keep a separate human listen. The notes say it plainly: a valid receipt proves submitted text, not pronunciation. A name read wrongly passes the receipt and fails the ear.
Why this is worth doing before the volume arrives
The cost of an audit trail is lowest on day one. If every job names a revision and sentence ids from the start, a later question like 'which narrations used the old price?' is a lookup. If jobs were sent as free text, the same question is a listening session over every file, and you pay for any regeneration on top. At $47.50 per 1M characters on Sume's router a 1,000 character retake is about five cents, but ten thousand of them is not.
Also decide who may change the script. The notes say agent credentials cannot accept or replace script bytes through the accept endpoint, and that a stale or other-thread revision fails closed. Treat that as the design: approval is a human act, generation is mechanical, and the receipt ties the two together.
Where the receipt sits in a release checklist
A practical checklist has four lines. One: approved script revision id recorded. Two: every voice job names that revision and its sentence ids. Three: the verify step returns full coverage in source order. Four: a person listened to each file and signed. The first three are automatable and cost nothing to read; the fourth is where you spend time.
Do not let a model's own claim, such as 'price lines exact', stand in for step three. The notes call out that matching sentence counts and model assertions cannot establish input integrity.
Sources
Related posts
More in Developers
- PWA manifest icons: any, maskable and monochrome from AI images
Make the three PWA icon purposes with the Sume image API: transparent, full-bleed maskable and one-color monochrome masters, plus the cost of each.
- Pydantic AI ToolCallJudge: check a Sume paid call before it runs
Pydantic AI 2.53.0 adds ToolCallJudge to assess tool calls before execution. Pair it with Sume's dry_run and max_spend_usd on paid calls.
- Pydantic AI cancel_and_resume vs Sume agent run cancel
Pydantic AI 2.54.0 shows cancel_and_resume. Sume Agent Completions have a different cancel: a run is cancelled by id and resumes only as a new run.
- Pydantic model to Sume output_schema: extra forbid, no defaults
Turn a Pydantic v2 model into a valid Sume Format output_schema: extra=forbid, nullable instead of defaults, and the SumeMediaFile reference. Tested.
Written by Sume