Voiceover plus a Lyria bed for a reel in one agent session over MCP
Ask an MCP client to run tts_create and music_create on Sume, then check the take with stt_create. Which tool does which job, and where the agent must wait.

In an MCP client connected to Sume, three tools cover a spoken reel: tts_create for the voice, music_create for the bed, and stt_create to check the take. All three are write tools, so a read-only session will refuse them. Run the voice first, then the music, then the check, and wait for each job to finish before the next.
The tools and their gates
The MCP docs list music_create, tts_create and stt_create among the generation tools that need write access. Two read tools sit beside them for scripts: tts_source_get, which returns the accepted-script manifest, and tts_source_verify_spine, which compares the selected TTS jobs with that script. Reads are free.
| Tool | Kind | Use in a reel |
|---|---|---|
| tts_create | Write | Voiceover from a transcript or an accepted script source |
| music_create | Write | Instrumental bed from a brief |
| stt_create | Write | Transcribe the take to check wording and timings |
| tts_source_get | Read | Fetch sentence ids of the accepted script |
| tts_source_verify_spine | Read | Check jobs cover the script |
A prompt that gets the order right
Agents do better with explicit sequencing and a stop rule. Something like the text below keeps the client from firing all three at once or resubmitting after a timeout.
Make a 20-second reel soundtrack.
1. Preview the cost of tts_create with dry_run, then run it for this script in English. Wait until the job is complete.
2. Read the word timings. Write a music brief with time-range markers that match the sentences, ending in "Instrumental, no vocals, no spoken word". Run music_create once.
3. Run stt_create on the voice file and compare the text to the script.
If a call times out, poll the job. Do not submit a second job.Where the agent must wait
The check step is useful because a valid receipt on a source-bound TTS job proves the submitted text, not the pronunciation, so a listen or an STT pass still earns its place.
- Sync waits end at 30 seconds. A timed-out wait means the job is still running.
- Music takes a prompt of 1 to 5,000 characters and no
durationfield. The agent must carry the length in the words. - TTS needs a voice selector. If the agent has only an API key, listing avatars and picking one with
voice.statusready is the discoverable route. - Raise a cost cap before a loop of takes, and keep one idempotency key per intended take.
Sources
Related posts
More in Integrations
- How to add an MCP server to ChatGPT with developer mode
Turn on ChatGPT developer mode, create an app for the server's URL, and sign in with OAuth. The steps, with Sume's hosted MCP server as the example.
- How to add subtitles to a video in Python
Add subtitles to a video in Python with Requests: POST the video URL to Sume's /v1/video-captions, poll the job, then read the captioned video_url.
- Add Sume to Claude as a custom connector (remote MCP)
Add Sume's hosted MCP server to Claude under Customize > Connectors, see what Sume's OAuth consent grants, and decide whether to allow paid tools.
- Airflow HTTP sensor: wait for an AI video job to finish
Submit an AI video job with Airflow's HttpOperator, then wait with an HttpSensor in reschedule mode that passes once the job's status is completed.
Written by Sume