Can ChatGPT make videos with sound? Audio and voice

Yes, through a video tool whose model generates audio. A voice that speaks your script is a second step: text to speech, then lip sync.

5 min readSume
All posts

Yes, if ChatGPT calls a video tool whose model generates audio: some video models return a clip with its own soundtrack, others return silent video. A person speaking your exact script is a different job, done in two steps: text to speech for the voice, then a lip-sync model that animates a still to that audio.

Sora is not the route: OpenAI says the Sora web and app experiences were discontinued on April 26, 2026 (OpenAI Help Center). ChatGPT reaches a video tool through developer mode, which provides full MCP client support; How to add an MCP server to ChatGPT covers setup. The Sume facts below come from Video Generation, the Models page, and MCP tools and gates, read on 2026-09-29. Hosted MCP still works but is not part of Sume's primary path today (basics).

Will the video have sound by default?

On Sume's generate_video tool, yes when the model supports it. The request field generate_audio defaults to the model's audio capability, and in current code the tool's description tells ChatGPT to prefer generate_audio: true, noting that leaving it out is also audio-on. It sets false only when you ask for a silent clip.

Not every model makes audio. Each model's catalog row carries a generate_audio flag that says whether it can generate an audio track, and ChatGPT can read it with the video-router_models tool before it submits. AI video with sound lists which models make sound.

From Video Generation, the Models page, and the Sume API reference, read 2026-09-29.
The sound you wantTool ChatGPT callsWhat to know
Sound generated with the clipgenerate_videoOnly on models whose catalog generate_audio flag is true
A voice speaking your scripttts_create, then avatar-image-to-video_createA still plus the voice audio; video models don't lip-sync
A silent clipgenerate_video with generate_audio: falseOnly on models that allow it; in current code a model that always makes audio refuses false

Can ChatGPT make a video where someone speaks my script?

Yes, but not with a video model alone. Sume's docs say video models do not lip-sync to generated TTS or to a later voice-over, so laying narration under a generated face will not match the lips. The documented route is:

  • tts_create voices the script with a Sume voice and returns an audio file. ChatGPT text to speech covers voices and prompts.
  • avatar-image-to-video_create (VEED Fabric 1.0, veed/fabric-1.0) turns one still plus that audio into a talking clip.
  • The audio must be on the Sume media host and at most 10 MB, and duration_seconds runs from 1 to 300.
  • Send exactly one visual source: a public HTTPS image_url, or a ready avatar by id or handle.
  • Output is 480p or 720p, with 720p the default.

How much does a video with sound cost?

A generated clip is billed per job from your Sume workspace balance at the provider's list price × 1.25; ask ChatGPT to call the tool with dry_run=true first, which previews the cost without submitting. A speaking clip has two meters on API pricing: text to speech at $0.0475 per 1,000 characters, and VEED Fabric 1.0 at $0.1875 per audio second (720p). Each rate carries a 5.5% agent fee on top by default.

Every paid call needs an idempotency_key, and max_spend_usd caps a call only when you send it. ChatGPT asks you to confirm write actions by default before it runs them.

What doesn't this do?

  • It does not read a video or audio file from your computer: hosted MCP cannot read files from your laptop.
  • It does not lip-sync a text-to-video clip to your voice-over; that is the Fabric step above.
  • It does not promise how the generated sound will sound. Listen to the clip before you use it.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume