tts_source_get and verify_spine: check a voiceover vs its script

Sume's hosted MCP has two free read tools for voiceover work. One returns the accepted script for tts_create, the other compares chosen TTS jobs to it.

5 min readSume
All posts

When an agent builds a voiceover from many tts_create calls, it can check its own work with two free reads. The hosted tools page lists tts_source_get, described as the accepted-script manifest for tts_create with transcript_source, and tts_source_verify_spine, which compares the selected TTS jobs with the accepted script (read 2026-10-05 in the repo docs). Both are read tools, so they spend nothing.

What the docs say, and what they do not

The page gives one line for each tool and no parameter list, so this post does not invent one. Call tools_schema with each tool's name to read the live contract before you use it. What is documented is the role: one tool shows what script was accepted, the other checks the jobs against it.

That is useful because a voiceover assembled from separate jobs can drift. A sentence could be generated twice, skipped, or taken from an older take. A check against the accepted script catches that before the audio goes into a timeline.

Voiceover tools on hosted MCP, read 2026-10-05
ToolKindRole
tts_createPaid, needs idempotency_keyCreate speech; routes to sume/auto unless a family is named
tts_source_getRead, freeAccepted-script manifest for a transcript_source create
tts_source_verify_spineRead, freeCompare selected TTS jobs with the accepted script
jobs_waitReadWait on one job, or up to 20 ids
script_runProgrammaticLoop many tool calls in one program

Fan out safely, then verify

The tools page suggests script_run when a turn needs three or more independent calls of the same shape, such as one tts_create per sentence. Inside the script, sume.call has the same gates as a direct call, so each paid create still needs its own idempotency_key. The script has timeout_seconds between 5 and 55, plus max_calls and max_paid_calls.

After the jobs finish, call tts_source_verify_spine on the selected jobs. If it reports a mismatch, regenerate only the affected sentence, with a new key, and verify again.

A local sanity check

Before you call the Sume verifier, a cheap local check catches an obvious count mismatch. This helper compares sentence count to job count and runs offline. It does not replace the Sume tool.

import re

def sentences(script: str) -> list:
    parts = re.split(r"(?<=[.!?])\s+", script.strip())
    return [p for p in parts if p]

def count_ok(script: str, job_ids: list) -> bool:
    return len(sentences(script)) == len(job_ids)

if __name__ == "__main__":
    script = "Meet the new lamp. It folds flat. Order today."
    print(count_ok(script, ["j1", "j2", "j3"]))
    print(count_ok(script, ["j1", "j2"]))

Where the cost sits

The two read tools cost nothing, but the creates they check do. Each tts_create is a paid call that needs idempotency_key, and max_spend_usd applies only if you send it. Preview with dry_run=true for the first sentence, then let the script spend under max_paid_calls.

A verification failure on one sentence should cost one regeneration, not the whole set. That is the practical reason to keep one job per sentence.

Checklist

For an agent that assembles speech from several jobs, adopt this order.

  • Fetch the accepted script with tts_source_get.
  • Create one job per sentence under max_paid_calls.
  • Wait with a batch jobs_wait, never by resubmitting.
  • Verify with tts_source_verify_spine before building the timeline.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume