Can ChatGPT make slideshow videos? Stills, voice, one MP4
Yes, through tools: ChatGPT can make the slides, voice the script, and have a timeline tool hold each still for its line in one MP4.

Yes, through tools: ChatGPT can write the narration, have each slide made as an image, voice the script, and have a timeline tool hold each still on screen for its line, then render the whole slideshow as one MP4. The pictures have to be files the video service already hosts, such as stills ChatGPT generated there, not photos on your computer.
ChatGPT calls those tools through developer mode, which gives it full MCP client support; How to add an MCP server to ChatGPT covers setup. The render facts come from Sume's Timeline 1.0 docs and MCP tools and gates, read on 2026-09-29. Sume has no official ChatGPT connector, and its basics page says hosted MCP still works but is not part of the primary path today.
How does ChatGPT turn pictures into a slideshow video?
ChatGPT writes one line of narration per slide, then works through Sume's hosted MCP tools. For three or more calls of the same shape, such as one generate_image per slide, script_run can send them in one tool call with the same gates as direct calls.
- Each still is a static hold: a
motionfield is accepted and ignored, so slides don't zoom or pan. - The sentence timings give each slide its
startandduration. Slideshow with AI voiceover shows how the numbers map. - For music,
music_createmakes a track and the timeline's optionalsoundtracklays it under the voice. A music-only slideshow setsaudio.modeto"silence"; Image slideshow video API covers that variant and the fit and transition options.
| Step | Tool | What comes back |
|---|---|---|
| Make each slide | generate_image | A Sume-hosted still |
| Voice the narration | tts_create | The audio file, its length, and, when ChatGPT asks for them, word and sentence timings |
| Hold each still for its line | timeline_create, then jobs_wait and timeline_get | One MP4, 1080×1920 by default, up to 1,800 seconds and 200 slots |
Can I use my own photos?
Not directly. Every URL in a timeline must already be your workspace's media.sume.com artifact or asset, and off-host URLs such as https://example.com/… are refused. Hosted MCP cannot read files from your laptop either. So the slides are stills Sume generated, for example with generate_image.
How much does a slideshow video cost?
The render is listed at $0.10 per output minute on API pricing, and the reserve is ceil(audio.duration_seconds / 60) minutes. Narration is $0.0475 per 1,000 characters. Image and music calls are billed per call from your Sume wallet, and every rate carries a 5.5% agent fee on top by default. Ask ChatGPT to run each paid tool with dry_run=true first, which previews the cost without submitting; ChatGPT also asks you to confirm write actions by default.
Sources
Related posts
More in Agents
- Can ChatGPT make TikTok videos? Vertical clips and limits
ChatGPT can write the hook and, with a video tool over MCP, have a 9:16 clip rendered. Posting it to TikTok is a separate step you do.
- Can ChatGPT make UGC videos? Script, voice, and lip sync
Not by itself. With a video tool connected in developer mode, ChatGPT can script, voice, and lip-sync a UGC-style clip, confirming paid calls first.
- Can ChatGPT make videos for free? Plans and per-clip costs
Not for free. Free ChatGPT accounts can't turn on developer mode, and the video tool it calls needs its own paid plan and bills each clip.
- Can ChatGPT make videos from photos? Image to video now
Not with Sora, which OpenAI discontinued. ChatGPT can send a photo's public URL to an image-to-video tool over MCP and get a clip back.
Written by Sume