After Sora: keep the recipe, not the prompt, as your stable layer
OpenAI named no replacement for the Sora Videos API. Put house style and output rules in a Sume Format so the next model change does not rewrite your code.

If the Sora shutdown taught you anything, it is that a prompt tied to one model is the most fragile thing in a video pipeline. The better layer to keep is the recipe: house style, output contract and quality bar, stored once and reused. A Sume Format stores that recipe, so your backend sends a small input and the model underneath can change.
What actually ended
OpenAI's deprecations page lists the Videos API and five Sora model ids as deprecated, with removal on September 24, 2026, announced March 24, 2026. The replacement column names nothing. That leaves each team to choose its own successor.
| Item | Status |
|---|---|
| Videos API | Deprecated, removal September 24, 2026 |
| sora-2, sora-2-pro | Deprecated |
| sora-2-2025-10-06, sora-2-2025-12-08, sora-2-pro-2025-10-06 | Deprecated |
| Recommended replacement | None listed |
Where the prompt fails you
A prompt that reads well for one model often needs rewriting for the next: different length limits, different handling of references, different audio behavior. If each call site builds its own prompt string, every call site needs the rewrite. If the prompt lives in a single recipe, one edit fixes all callers.
What a Format holds
A Format is a saved production recipe: a house style, an output contract and a playbook for one kind of video. Your backend calls it by name with POST /v1/formats/{handle}/{slug}/runs. The recipe is set before your instruction, so you do not resend a system prompt on every call. If your instruction and the recipe disagree, the model follows your instruction (Format API).
- Style and rules live in the Format, versioned. Each receipt shows which
format.versionran. - Your per-call data lives in
input: product URL, script, host image. - You can bind a JSON Schema so the result comes back in your own shape (Structured output).
- A spend cap per run limits the cost of a bad recipe.
When a Format is the wrong layer
A Format run is an agent turn that takes minutes, not a single model call. If you only need one clip from a prompt, POST /v1/videos with a model id or sume/auto is simpler (Generate videos). Use a Format when the job has several steps, such as host takes, B-roll, voiceover and assembly, and when you want the same look every time.
Moving to a Format also does not make output identical to what Sora produced. Expect to re-judge quality on a sample set, and keep the old files as a reference.
A migration order
First move the call sites that share a style into one Format. Second, test it against ten of your old prompts. Third, set a spend cap and a webhook, then switch traffic. The next time a model is retired, you edit the recipe once and read the receipts.
Sources
Related posts
More in Formats
- Bulk queue concurrency 16 on Sume is a window, not a speed promise
Sume bulk Format runs accept concurrency 1 to 16. It limits how many children start at once and does not promise throughput. What 100 ad items really do.
- 100 Omni proofs, then 20 finals: two bulk queues cost $45, not $112.50
Run 100 weekly ad proofs at 360p in one Sume bulk queue, approve 20, then queue those at 1080p. Video cost: $22.50 plus $22.50 instead of $112.50.
- Format showcase field: judge a Format by a real output first
GET /v1/formats returns showcase, a verified sample output. Use it with description and io to pick a Sume Format before you spend credits on a run.
- Format spend caps: the $400 default, $500 maximum, and your own
A Sume Format run can never spend past its cap. How the cap is chosen, what happens when a run hits it, and how to size generation_spend_cap_usd per run.
Written by Sume