Sume TTS transcript_receipt: it proves the text, not the pronunciation

A source-bound Sume TTS job returns a transcript_receipt with a SHA-256 of the submitted text. It proves what was sent, not how it was spoken. Listen too.

4 min readSume
All posts

A TTS job on Sume that was submitted from a source revision carries a server-owned transcript_receipt. It records the job id, the revision, the sentence ids, the canonicalization version, the SHA-256 of the submitted transcript, and input_integrity set to source_bound. That proves which text went in. It says nothing about how the voice read it, so you still need to listen.

What the receipt holds

The receipt is written by the server, not by the caller, and is still present when you read the job back after submit, status and result. Caller metadata cannot set it or certify it.

Sume TTS receipt facts, read 2026-10-06 from the repo docs
Receipt fieldWhat it tells you
job idWhich job the receipt belongs to
revision and sentence idsWhich accepted script text, in which order
canonicalization versionHow the text was normalised before hashing
submitted transcript SHA-256Exact bytes that were sent
input_integrity: source_boundThe text came from a persisted source, not literal input

What it does not tell you

A matching hash means the text was right. A voice can still mispronounce a name, flatten a question, or read 2026 as a number rather than a year. Sentence counts that match, or a model's claim that prices are exact, do not prove integrity either, which is why the receipt exists. Keep the two checks apart: the receipt for the text, a listening pass for the speech.

  • Receipt: was the right text submitted?
  • Listening: was it spoken correctly?
  • STT round trip: did every word come back?

Using it in a release check

A simple gate has two lines. First, compare the receipt's revision and sentence ids with the plan: every sentence in the approved script must appear once, in order, across the accepted jobs. Second, have a person listen to the finished audio, or run a speech-to-text round trip and compare it against the script.

If the first check fails, reject the batch. Do not repair it by hand or waive it, since the point of source binding is that nobody can talk the pipeline past a mismatch.

When it matters

Source binding is built for regulated or price-sensitive reads, such as live commerce price lines, where the spoken text must match an approved script. For an ordinary voiceover, a literal transcript works. A missing receipt on an old job should not trigger regeneration on its own; the verify route can adopt an exact text match without a new provider call.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume