Can ChatGPT make videos from photos? Image to video now

Not with Sora, which OpenAI discontinued. ChatGPT can send a photo's public URL to an image-to-video tool over MCP and get a clip back.

5 min readSume
All posts

Not with Sora anymore, but yes with a video tool: ChatGPT can turn a photo into a short video clip when you connect an image-to-video tool in developer mode and give it the photo's public URL. OpenAI says the Sora web and app experiences were discontinued on April 26, 2026, and its help page gives September 24, 2026 as the Sora API's discontinuation date.

OpenAI's side comes from its Sora discontinuation article and ChatGPT developer mode guide. Sume's side comes from Video Generation and MCP tools and gates. All were read on 2026-09-29. Sume has no official ChatGPT connector: this is a developer-mode app pointed at Sume's hosted MCP server, and Sume's basics page says hosted MCP still works but is not part of the primary path today. For the plain API version, read How to make a picture move with AI.

How do I get ChatGPT to turn my photo into a video?

Connect a video tool in developer mode, then ask for it by name, as OpenAI suggests. Can ChatGPT make videos? walks through the developer-mode setup, the Write toggle on Sume's consent page, and the approval prompt. With Sume's hosted MCP server the tool is generate_video, and the photo goes in as its first frame:

Use the Sume app's generate_video tool. Animate the photo at
https://example.com/photos/lake.jpg as the first frame: slow
push-in, water rippling. 9:16, 5 seconds. Run dry_run first
and show me the cost.

Does my photo need to be online?

Yes. Sume's hosted MCP server cannot read files from your laptop, and its video docs say reference images must be accessible over public HTTPS. Put the photo at a public HTTPS URL first, for example on your own site or storage, and paste that URL into the chat. In current code, a reference URL that answers 404 fails the job with input_media_unreachable.

When the job finishes, the clip comes back as a Sume-hosted artifact under media.sume.com, and ChatGPT can hand you that link.

Should the photo be the first frame or a reference?

It depends on how closely the video should match the photo. The request has two image fields, and each one switches the model into a different mode.

From Video Generation, read 2026-09-29.
FieldModeWhat the photo does
frame_images with first_frameImage-to-videoThe clip starts on your photo
frame_images with last_frameImage-to-videoThe clip ends on your photo; current code refuses it without a first_frame
input_referencesReference-to-videoVisual guidance for style or content, not an exact frame
Both fieldsImage-to-videoframe_images takes precedence

How long can the clip be, and what does it cost?

Limits differ per model. Sume's docs list seedance-2.5 at 4–30 seconds and wan-3.0 at 2–30 seconds, and say every other catalog model tops out at 15 seconds. ChatGPT can read each model's supported_durations, supported_resolutions, and supported_frame_images with the video-router_models tool before it submits.

Billing is per job from your Sume workspace balance, reserved on submit at the provider's list price × 1.25, plus a 5.5% agent fee by default. Ask ChatGPT for a dry_run=true call first: it previews the cost without submitting the job.

A video longer than one clip means joining several; see Can ChatGPT make long videos?. Can ChatGPT make videos? covers Sora's shutdown and how to wait for a finished job.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume