Sora API ended: which Sume route replaces a talking presenter clip?

OpenAI's Sora API ended 2026-09-24. For a presenter who says exact words, the replacement is a still-plus-audio or avatar route, not a video model.

5 min readSume
All posts

If you used the Sora API to produce a person speaking, the closest Sume replacement depends on whether you need exact words. For a presenter who must say your script, use Avatar 1.0 talking video (4 to 60 seconds) or the H3 Max lip-sync route (5 to 14.8 seconds); for a scene where a person speaks as part of a generated shot, the video models with native audio are the candidates. A tracker lists the OpenAI Sora API as ended on 2026-09-24 and Sora as discontinued (Magic Hour tracker, read 2026-10-06).

Do not assume a prompt-driven video model will lip-sync to your words. Sume's models docs state that video models do not lip-sync to generated TTS or to a later voice-over, so a face that talks is a still plus audio through a lip-sync model (Models, read 2026-10-06).

Which job maps to which route?

Routing by the job you used Sora for. Route facts from docs.sume.com (Models, Video generation, Generate avatar video), read 2026-10-06.
What you madeSume routeConstraint
A presenter reading your exact scriptPOST /v1/avatar-1.0/talking-videoReady avatar, 4 to 60 s, 720p
Your photo, your voice file, a short linePOST /v1/minimax/h3-max/lip-syncSume-hosted audio, 5 to 14.8 s
A scene with ambient sound and no required wordsPOST /v1/videos with an audio-capable modelPick the id in the video catalog
A character copying a move you filmedPOST /v1/kling/3.0/motion-controlMotion video you supply, 1 to 30 s declared

What changes in your code?

Sora's integration was a prompt in, video out. A talking clip on Sume is two inputs you control: who is on screen and what they say. Create an avatar once (prompt, profile or photo), keep its avatar_handle, and send a script. Aspect ratios are 1:1, 3:4, 9:16, 4:3 and 16:9, and captions can be burned inline. The job lifecycle is the same as every Sume generation: submit, receive a job id, poll or take a webhook, read the result, and store the Sume-hosted URL rather than a provider URL.

For the plain video-model side, the model-per-prompt post maps old prompts by clip length, and the Go client post shows a client that handles Sume jobs. Neither of those covers a speaking presenter, which is the gap this post is about.

What should you test first?

Take one script of 20 seconds you used to generate with Sora and run it three ways: an avatar talking video at standard, an H3 lip-sync clip at 480p with Sume TTS audio, and a native-audio video prompt. Compare how closely each says the words, whether the face stays consistent between shots, and what you pay. Expect the first two to match a script exactly and the third to take creative liberties.

Write down the one requirement that rules out the others. If it is word accuracy, you are choosing between the first two. If it is cinematic camera work, you are choosing a video model and accepting that the speech will not be your script. Mixed jobs often need both: a lip-sync talking head for the line and a video clip for the cutaway, joined on a timeline.

What should you migrate first?

Start with the clips that carried speech, because those are the ones a plain video model cannot replace. List them, note each line's length, and sort into those under 14.8 seconds, which suit H3 Max lip sync, and those up to 60 seconds, which suit the Avatar 1.0 talking-video route. Prompts for silent scenes can move to a catalog video model after a test render.

Keep your old prompts but do not expect identical output. A different model reads a prompt differently, so budget a review pass for each migrated clip and compare against the old render side by side. Keep the Sora API end date in the project notes, 2026-09-24 per the tracker read on 2026-10-06, so a reviewer understands why the endpoint is gone and not just broken.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume