Prove a voiceover read the approved script: the transcript receipt

A source-bound Sume TTS job carries a transcript receipt with the revision, sentence IDs and a SHA-256. It proves the text sent, not how it was pronounced.

6 min readSume
All posts

If a reviewer asks whether the voiceover says what legal approved, a Sume TTS job built from a source-bound script can answer for the text. The job exposes a server-owned transcript_receipt with the job id, script revision, sentence ids, canonicalization version, the SHA-256 of the submitted transcript and input_integrity: "source_bound". A read-only verify step then checks selected jobs against the accepted script. What the receipt cannot prove is how the words sound.

With new speech models arriving at $15 to $22 per 1M characters from Microsoft (read 2026-10-03), more teams will generate narration at volume. Volume makes an audit trail worth having.

What a receipt covers

Raw transcript text, paths, client hashes and receipt fields are rejected inside transcript_source. You name sentences by id and the server resolves them, which is why a model cannot quietly edit a price line on the way to the voice.

From Sume's source-bound TTS notes and the MCP tools page, read 2026-10-03.
QuestionAnswer
Text input ruleExactly one of transcript or transcript_source
transcript_source fieldsscript_revision_id and sentence_ids only
Discover the accepted scriptGET /v1/tts-1.0/source (tool tts_source_get, free)
Check jobs against itPOST /v1/tts-1.0/source/verify-spine (tool tts_source_verify_spine, free)
Receipt holdsJob id, revision, sentence ids, canonicalization version, transcript SHA-256, input_integrity
Receipt does not provePronunciation, pacing, acoustic quality
TTS RouterSame contract, plus the required model

A checker for your own records

Even without calling the verify route, you can keep your own ledger. Hash each approved sentence at approval time, hash what each job reports, and compare. The snippet below shows the shape of the comparison with local data, so you can adapt it to whatever your jobs return.

import hashlib

def sha(text):
    return hashlib.sha256(text.encode("utf-8")).hexdigest()

approved = {
    "s1": "Order by October 31.",
    "s2": "Free shipping on orders over 50 dollars.",
}
# what you recorded for each generated job: sentence ids and the transcript hash
jobs = [
    {"job": "job_a", "sentence_ids": ["s1"], "transcript_sha256": sha("Order by October 31.")},
    {"job": "job_b", "sentence_ids": ["s2"], "transcript_sha256": sha("Free shipping on orders over 40 dollars.")},
]
for j in jobs:
    expected = sha(" ".join(approved[i] for i in j["sentence_ids"]))
    ok = expected == j["transcript_sha256"]
    print(j["job"], "MATCH" if ok else "MISMATCH")

What to do with a mismatch, and with a missing receipt

Per the notes, mismatched text cannot be adopted, and a missing receipt on its own must never trigger automatic regeneration. For older jobs with no receipt, you name the sentence ids and the server compares the stored submitted text exactly after canonicalization; an exact match can be adopted with no writes, provider calls or regeneration. That is the cheap path, so try it before you pay to regenerate.

Keep a separate human listen. The notes say it plainly: a valid receipt proves submitted text, not pronunciation. A name read wrongly passes the receipt and fails the ear.

Why this is worth doing before the volume arrives

The cost of an audit trail is lowest on day one. If every job names a revision and sentence ids from the start, a later question like 'which narrations used the old price?' is a lookup. If jobs were sent as free text, the same question is a listening session over every file, and you pay for any regeneration on top. At $47.50 per 1M characters on Sume's router a 1,000 character retake is about five cents, but ten thousand of them is not.

Also decide who may change the script. The notes say agent credentials cannot accept or replace script bytes through the accept endpoint, and that a stale or other-thread revision fails closed. Treat that as the design: approval is a human act, generation is mechanical, and the receipt ties the two together.

Where the receipt sits in a release checklist

A practical checklist has four lines. One: approved script revision id recorded. Two: every voice job names that revision and its sentence ids. Three: the verify step returns full coverage in source order. Four: a person listened to each file and signed. The first three are automatable and cost nothing to read; the fourth is where you spend time.

Do not let a model's own claim, such as 'price lines exact', stand in for step three. The notes call out that matching sentence counts and model assertions cannot establish input integrity.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume