ChatGPT script to video: voice it, then cut to each line
ChatGPT writes the script; tools turn it into video. Voice it word for word, make one visual per line, and cut each to its sentence on a timeline.

ChatGPT can write the script, but turning that exact script into a video takes tools. In developer mode ChatGPT can voice the script word for word with a text-to-speech tool, make one clip or still per line, and cut each visual to its sentence on a timeline that renders one MP4. Without those tools, what you get from the chat is the script.
ChatGPT's side comes from OpenAI's ChatGPT Developer mode guide. The tools are on Sume's hosted MCP server, per MCP tools and gates, Timeline 1.0, and the Sume API reference, read on 2026-09-29. Sume has no official ChatGPT connector: this is a remote MCP connection, and Sume's basics page says hosted MCP still works but is not the primary path today. The general method, without ChatGPT, is in Script to video AI.
How does ChatGPT turn a script into a video?
In three tool steps after the script is final. Once developer mode is on (OpenAI lists it for Pro, Plus, Business, Enterprise, and Education accounts on the web) and Sume's app has Write turned on, ChatGPT calls:
tts_createwith the script astranscript,timestamps.words: true, andsegmentation.mode: "sentence". The finished job then carrieswords[]with start and end seconds and gapless sentencesegments[], so each sentence has a start and an end.generate_videoorgenerate_imageonce per line, for a clip or a still that shows it.timeline_createwith the voiceover as the audio spine and onevideo[]slot per line, each starting where its sentence starts. Thenjobs_waitandtimeline_getfor the MP4.
Can ChatGPT make all the per-line clips in one call?
Yes, with script_run. It runs a short JavaScript program on Sume's side for turns that need three or more calls of the same shape, such as one generate_image per scene. Each call inside keeps the same gates, and each paid create still needs its own idempotency_key. The run is bounded by timeout_seconds (5–55), max_calls, and max_paid_calls, and it returns the child jobs[] to wait on, so ChatGPT follows it with jobs_wait rather than expecting finished clips.
Will ChatGPT keep every word of my script?
Only if the words it sends are yours. tts_create speaks the transcript it receives, and ChatGPT fills that field, so freeze the script first and read the tool input before you approve. OpenAI says the full JSON of each tool call's input is available in the chat and asks you to review write actions carefully. The checks worth making at each step:
- The voiceover, clips, and stills from the earlier steps are Sume-hosted, so they qualify. A file from your computer does not: hosted MCP cannot read your disk.
- In current code a Timeline render takes sound only from the voiceover spine and an optional soundtrack; each clip's own audio is dropped. Script to video AI covers the length and slot limits.
| Tool call | Check in the input | Why |
|---|---|---|
tts_create | transcript matches your script word for word | Up to 20,000 characters are voiced as sent |
tts_create | timestamps.words: true and segmentation.mode: "sentence" | Sentence segments[] are the cut points |
| Clips and stills | dry_run: true on the first call | Previews the cost without submitting |
timeline_create | video[0].start is 0; each later start is a sentence start | Declared starts are authoritative |
timeline_create | Every URL is a media.sume.com file | Timeline takes only this workspace's Sume-hosted media |
How much does a script-to-video run cost?
The voiceover bills $0.0475 per 1,000 characters; the timeline render bills $0.10 per output minute, reserved per started minute; clips and stills are priced per model. Each paid call adds a 5.5% agent fee by default. OpenAI says write actions require confirmation by default, and Sume's dry_run=true previews a paid call's cost without submitting it, so ask ChatGPT to dry-run the voiceover and the clips and show you the prices first. One jobs_wait holds at most 55 seconds; MCP tool call timeouts on long video jobs covers the wait loop.
Sources
Related posts
More in Agents
- Run the Sume video agent from your backend with Agent Completions
POST /v1/agent/completions runs the same agent as the Sume Agents chat, with tools and media generation, and returns an async run receipt you poll or webhook.
- Safe automation for AI agents that call paid APIs
Keep agents read-only by default, keep secrets out of logs, and on hosted MCP send an idempotency_key, preview with dry_run, and cap with max_spend_usd.
- Scheduled AI video agent runs: cron, API triggers, and receipts
A Sume schedule is a saved Agents automation that runs on a cron cadence and returns a run receipt. Author it in the dashboard; start and monitor runs by API.
- What is a video agent? How Sume defines and runs one
In Sume's docs, a video agent is a sandbox Agent that composes generation tools into a post-ready video. Brief it in chat, or call it over HTTP.
Written by Sume