Voiceover plus a Lyria bed for a reel in one agent session over MCP

Ask an MCP client to run tts_create and music_create on Sume, then check the take with stt_create. Which tool does which job, and where the agent must wait.

5 min readSume
All posts

In an MCP client connected to Sume, three tools cover a spoken reel: tts_create for the voice, music_create for the bed, and stt_create to check the take. All three are write tools, so a read-only session will refuse them. Run the voice first, then the music, then the check, and wait for each job to finish before the next.

The tools and their gates

The MCP docs list music_create, tts_create and stt_create among the generation tools that need write access. Two read tools sit beside them for scripts: tts_source_get, which returns the accepted-script manifest, and tts_source_verify_spine, which compares the selected TTS jobs with that script. Reads are free.

Sume MCP audio tools - from the MCP tools and gates page (read 2026-10-07)
ToolKindUse in a reel
tts_createWriteVoiceover from a transcript or an accepted script source
music_createWriteInstrumental bed from a brief
stt_createWriteTranscribe the take to check wording and timings
tts_source_getReadFetch sentence ids of the accepted script
tts_source_verify_spineReadCheck jobs cover the script

A prompt that gets the order right

Agents do better with explicit sequencing and a stop rule. Something like the text below keeps the client from firing all three at once or resubmitting after a timeout.

Make a 20-second reel soundtrack.
1. Preview the cost of tts_create with dry_run, then run it for this script in English. Wait until the job is complete.
2. Read the word timings. Write a music brief with time-range markers that match the sentences, ending in "Instrumental, no vocals, no spoken word". Run music_create once.
3. Run stt_create on the voice file and compare the text to the script.
If a call times out, poll the job. Do not submit a second job.

Where the agent must wait

The check step is useful because a valid receipt on a source-bound TTS job proves the submitted text, not the pronunciation, so a listen or an STT pass still earns its place.

  • Sync waits end at 30 seconds. A timed-out wait means the job is still running.
  • Music takes a prompt of 1 to 5,000 characters and no duration field. The agent must carry the length in the words.
  • TTS needs a voice selector. If the agent has only an API key, listing avatars and picking one with voice.status ready is the discoverable route.
  • Raise a cost cap before a loop of takes, and keep one idempotency key per intended take.

Sources

Related posts

More in Integrations

All Integrations posts

Written by Sume