Sora API ended: which Sume route replaces a talking presenter clip?
OpenAI's Sora API ended 2026-09-24. For a presenter who says exact words, the replacement is a still-plus-audio or avatar route, not a video model.

If you used the Sora API to produce a person speaking, the closest Sume replacement depends on whether you need exact words. For a presenter who must say your script, use Avatar 1.0 talking video (4 to 60 seconds) or the H3 Max lip-sync route (5 to 14.8 seconds); for a scene where a person speaks as part of a generated shot, the video models with native audio are the candidates. A tracker lists the OpenAI Sora API as ended on 2026-09-24 and Sora as discontinued (Magic Hour tracker, read 2026-10-06).
Do not assume a prompt-driven video model will lip-sync to your words. Sume's models docs state that video models do not lip-sync to generated TTS or to a later voice-over, so a face that talks is a still plus audio through a lip-sync model (Models, read 2026-10-06).
Which job maps to which route?
| What you made | Sume route | Constraint |
|---|---|---|
| A presenter reading your exact script | POST /v1/avatar-1.0/talking-video | Ready avatar, 4 to 60 s, 720p |
| Your photo, your voice file, a short line | POST /v1/minimax/h3-max/lip-sync | Sume-hosted audio, 5 to 14.8 s |
| A scene with ambient sound and no required words | POST /v1/videos with an audio-capable model | Pick the id in the video catalog |
| A character copying a move you filmed | POST /v1/kling/3.0/motion-control | Motion video you supply, 1 to 30 s declared |
What changes in your code?
Sora's integration was a prompt in, video out. A talking clip on Sume is two inputs you control: who is on screen and what they say. Create an avatar once (prompt, profile or photo), keep its avatar_handle, and send a script. Aspect ratios are 1:1, 3:4, 9:16, 4:3 and 16:9, and captions can be burned inline. The job lifecycle is the same as every Sume generation: submit, receive a job id, poll or take a webhook, read the result, and store the Sume-hosted URL rather than a provider URL.
For the plain video-model side, the model-per-prompt post maps old prompts by clip length, and the Go client post shows a client that handles Sume jobs. Neither of those covers a speaking presenter, which is the gap this post is about.
What should you test first?
Take one script of 20 seconds you used to generate with Sora and run it three ways: an avatar talking video at standard, an H3 lip-sync clip at 480p with Sume TTS audio, and a native-audio video prompt. Compare how closely each says the words, whether the face stays consistent between shots, and what you pay. Expect the first two to match a script exactly and the third to take creative liberties.
Write down the one requirement that rules out the others. If it is word accuracy, you are choosing between the first two. If it is cinematic camera work, you are choosing a video model and accepting that the speech will not be your script. Mixed jobs often need both: a lip-sync talking head for the line and a video clip for the cutaway, joined on a timeline.
What should you migrate first?
Start with the clips that carried speech, because those are the ones a plain video model cannot replace. List them, note each line's length, and sort into those under 14.8 seconds, which suit H3 Max lip sync, and those up to 60 seconds, which suit the Avatar 1.0 talking-video route. Prompts for silent scenes can move to a catalog video model after a test render.
Keep your old prompts but do not expect identical output. A different model reads a prompt differently, so budget a review pass for each migrated clip and compare against the old render side by side. Keep the Sora API end date in the project notes, 2026-09-24 per the tracker read on 2026-10-06, so a reviewer understands why the endpoint is gone and not just broken.
Sources
Related posts
More in Comparisons
- Seedance 2, Omni or Wan 3.0 for 9:16 ads on Sume: how to choose
Three Sume video models side by side for vertical ads after Sora: clip length, resolutions, 9:16 in the docs, references, and the price of a 5 s clip.
- Sync.so $0.05 a second vs Sume H3 Max lip sync: rates side by side
Sync.so lists $0.05 to $0.04 a second on four plans; Sume H3 Max lip sync derives $0.0625, $0.10 and $0.20. What each rate buys.
- Synthesia Starter's 12 minutes vs twelve 60-second Sume avatar jobs
Synthesia lists $29 for about 12 minutes and $89 for about 60. How those minutes map to Sume avatar jobs capped at 60 seconds, and what each setup is for.
- Tavus Starter vs Growth: break-even is about 1,014 minutes a month
Tavus Starter is $59 plus $0.37 a minute over 100; Growth is $397 with 1,250 minutes. The crossover, and when a rendered Sume clip fits better than minutes.
Written by Sume