ChatGPT text to speech: make an audio file with MCP
ChatGPT Voice speaks its replies. To get your own script as an audio file, add a text-to-speech tool to ChatGPT over MCP in developer mode.

ChatGPT speaks its replies in a conversation: ChatGPT Voice lets you talk with ChatGPT and hear a spoken response, OpenAI's help page says. To turn your own script into an audio file you can download and reuse, give ChatGPT a text-to-speech tool: turn on developer mode, add a remote MCP server that has one, and ask ChatGPT to voice the script. With Sume's hosted MCP server at https://mcp.sume.com/mcp, that tool is tts_create, and ChatGPT returns a link to the finished file, an MP3 unless you ask for WAV.
ChatGPT's side comes from OpenAI's ChatGPT Voice help page and ChatGPT Developer mode guide. Sume's side comes from MCP tools and gates, MCP OAuth and API keys, and the TTS 1.0 schema in the Sume API reference. All were read on 2026-09-28. Sume's basics page says hosted MCP still works but is not the primary path today; a backend should call the text to speech API directly.
How do I connect a text-to-speech tool to ChatGPT?
Through developer mode, which OpenAI says provides "full Model Context Protocol (MCP) client support for all tools, both read and write" and labels Elevated risk. It is available to Pro, Plus, Business, Enterprise, and Education accounts on the web. Turn it on under Settings → Security and login, create a developer-mode app for https://mcp.sume.com/mcp with OAuth, and turn Write on at Sume's consent page, where it is off by default: a read-only session gets insufficient_scope from paid tools such as tts_create. Sume has no official ChatGPT connector; How to add an MCP server to ChatGPT walks through each screen.
What do I ask ChatGPT to make the audio file?
Choose Developer mode from the Plus menu and select the Sume app for the conversation. OpenAI suggests being explicit: name the app and the tool, and rule out other tools so ChatGPT doesn't reach for a different one. For example:
Use the Sume app's tts_create tool to voice the script below
with my narrator avatar. Only use the Sume app.
Run it with dry_run first and show me the cost.
Welcome to the tour. First, open the dashboard.Which voice does ChatGPT use?
The Sume voice you name: tts_create needs one, and the voice you can discover is a Sume avatar whose voice is ready. ChatGPT can find yours with the read-only avatars_list tool; voice.id works instead if you already hold a Sume voice id. A voice name from another service is rejected, so ChatGPT Voice's own voices, such as Juniper, don't carry over. Set language for any script that isn't English, because an omitted language defaults to English. Text to speech API lists every request field.
Which tool calls need my approval?
The tts_create calls, dry run included. OpenAI says write actions require confirmation by default, and ChatGPT respects the readOnlyHint tool annotation. In current code Sume marks avatars_list, jobs_wait, and jobs_result read-only and tts_create a write, so by default ChatGPT asks before the dry run and again before the real submit. Check the input it shows you each time: the script, the voice, and dry_run.
| Call | What it does | Approval |
|---|---|---|
avatars_list | Lists your avatars, so ChatGPT can pick one whose voice is ready | Read-only |
tts_create with dry_run: true | Previews the cost; no job is submitted | Write: confirm by default |
tts_create | Submits the job with its required idempotency_key and returns a job id, not audio | Write: confirm by default |
jobs_wait | Holds at most 55 seconds per call; while the job runs, the next step is another wait on the same id, never a second submit | Read-only |
jobs_result | Returns the finished job with the file's audio_url (current code) | Read-only |
Where is the audio file when the job finishes?
In the jobs_result output. OpenAI's guide says the full JSON input and output of each tool call are available when you expand it, so the audio_url is there even if ChatGPT's reply leaves it out. It points at a public artifact on media.sume.com that you can open or download outside the chat.
What does it cost, and what are the limits?
tts_create bills $0.0475 per 1,000 characters, plus a 5.5% agent fee by default, and a dry_run previews the cost without submitting the job. One call takes up to 20,000 characters, and audio longer than 1,200 seconds fails with tts_duration_exceeded, with no credit captured. The file is not live speech: it exists once the job completes. The other rules are the same from any MCP client; see Claude text to speech.
Sources
Related posts
More in Agents
- HeyGen Video Agent: what it does, its API, and pricing
HeyGen's Video Agent turns a text prompt into a finished avatar video, in the app or via POST /v3/video-agents. Its inputs, outputs, and API price.
- How to write a video brief: parts, template, agent tips
A video brief states the deliverable, audience, one message, must-show items, tone, and what done means. A template, and how an AI agent reads it.
- Image generation MCP server: how Sume's generate_image works
Sume's hosted MCP server has a paid generate_image tool: a prompt in, a job id back in milliseconds, then jobs_wait and jobs_result for the images.
- MCP vs function calling: how they differ and fit together
Function calling lets a model ask your app to run a function you defined; MCP puts tools on a server any client can discover. How they fit together.
Written by Sume