Put a specific person in AI video after the Sora API shutdown
The Sora Videos API shut down on September 24, 2026. To keep a specific person on screen, Sume has three routes: motion control, Recast and reference images.

The short answer
OpenAI's deprecations page lists the Videos API and the sora-2 models with a shutdown date of September 24, 2026. If your Sora workflow put a specific person on screen, Sume has three routes: Kling motion control at $0.1575 a second, H3 Max Recast at $0.375 a second at 768p, and reference images on MiniMax H3 or Gemini Omni Flash.
What changed
OpenAI's deprecations page says developers were notified on March 24, 2026, and that the Videos API and sora-2, sora-2-pro and their dated snapshots shut down on September 24, 2026. The page names no replacement. This post only covers the one job where the replacement choice is not obvious: keeping a particular person consistent.
Pick by where the movement comes from
The three routes differ in where the movement comes from.
| Route | Movement comes from | Person comes from | Billed rate |
|---|---|---|---|
| Kling motion control | A reference video | One photo | $0.1575 per second |
| H3 Max Recast | A source video you already have | 1 to 4 person photos | $0.375 per second at 768p, $0.5625 at 1080p |
| Reference images on H3 or Omni | The text prompt | Up to 9 (H3) or 10 (Omni) images | H3 768p $0.075 per second; Omni 720p $0.125 per second |
When each one fits
Choose motion control when you can film the move yourself. Choose Recast when the video already exists, such as a shot from a previous Sora run or a stock clip, and you want to swap the person in it. The source must run 5 to 30 seconds with no single shot longer than 15 seconds. Choose reference images when you want the model to invent the scene and only need the person to look right.
Reference-image request
Reference images ride in input_references on /v1/videos. The same call works for either model; only model changes.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: person-001" \
-d '{
"model": "minimax-h3",
"prompt": "The person in the reference image walks through a market",
"duration": 8,
"resolution": "768p",
"input_references": [
{"type": "image_url", "image_url": {"url": "https://example.com/person.png"}}
]
}'Consent
Use photos of people who agreed to appear in the video. For a migration of the plain submit-and-poll flow, see the Sora smoke test.
Sources
Related posts
More in Comparisons
- Luma Ray 3.2 features and the Sume endpoint for each one
Sume does not run Ray 3.2. Match its keyframes, reframe, face tracking and 20 s clips to the Sume surfaces that exist, with docs-verified limits.
- Ray 3.2 takes 16 keyframes; Sume frame_images takes two
Luma lists up to 16 keyframes per Ray 3.2 clip. Sume's frame_images array uses first_frame and last_frame. Which catalog models list them, per the docs.
- Ray 3.2 tracks 8 faces; Sume swaps 1-4 people with H3 Max Recast
Luma Ray 3.2 lists facial tracking for up to 8 faces. Sume has no tracking output, but h3-max-recast swaps 1-4 people in a clip and Kling drives a still.
- Reference limits per request, vendor vs Sume: Wan 3.0, Seedance, Omni
What each vendor says one video request can take, next to what Sume's catalog allows: Wan 3.0, Seedance 2.5, Gemini Omni 1.1 Flash, and Ray3.2.
Written by Sume