Sora API removed, no replacement listed: pick a Sume video id

OpenAI removed the Videos API on 2026-09-24 and lists no replacement. Pick a Sume video model by duration cap and the price of an 8-second 720p clip.

5 min readSume
All posts

OpenAI's deprecations page says the Videos API and the Sora 2 model aliases and snapshots were removed from the API on September 24, 2026, after an announcement on March 24, 2026, and its replacement column is empty. If you had a Sora 2 call in production, you need a different provider, and on Sume the cheapest 8-second 720p options are wan-3.0 and gemini-omni-flash-1.1 at $1.00 each, with seedance-2.5 at $4.6224.

This post does not compare output quality, because no source here supports a quality ranking. It sorts the Sume rows by the two things a port actually hits first: the duration cap and the price.

Pick by the cap you need

Sume documents the limits per model. wan-3.0 accepts 2 to 30 seconds, seedance-2.5 accepts 4 to 30 seconds at 480p, 720p and 1080p, and the other catalog rows have a 15-second ceiling unless noted. gemini-omni-flash-1.1 accepts 3 to 10 seconds, 16:9 or 9:16, with native audio always on. minimax-h3 accepts 5 to 15 seconds at native 480p or 768p, so it has no 720p tier at all.

Eight-second clip, Sume list price (as of 2026-10-09). Each row is the per-second price times 8.
Sume model idValid durationTier usedPer second8-second clip
wan-3.02-30 s720p$0.125$1.00
gemini-omni-flash-1.13-10 s720p$0.125$1.00
minimax-h35-15 s768p$0.075$0.60
minimax-h3-max5-15 s768p$0.10$0.80
seedance-24-15 s720p$0.378$3.024
seedance-2.54-30 s720p$0.5778$4.6224

Porting the request

The Sume Videos route mirrors the OpenRouter video shape, with bare catalog ids instead of provider-prefixed ones. For a direct port of a 16:9 prompt with a duration field, the Video Router is the simplest call. Send an Idempotency-Key so that a client retry does not queue a second clip.

curl -X POST https://api.sume.com/v1/video-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: sora-port-001" \
  -d '{
    "model": "wan-3.0",
    "prompt": "A vertical product clip on a desk, natural light",
    "resolution": "720p",
    "duration": 8,
    "aspect_ratio": "9:16",
    "mode": "async"
  }'

What does not carry over

The table above does not make any claim about Sora's lengths, resolutions or prompt behavior, because the only Sora source this post reads is OpenAI's deprecation entry. Two practical points follow from Sume's own docs. First, every Sume model has its own envelope, so read capabilities from the models endpoint before you hard-code a duration. Second, minimax-h3 is native 768p and not 720p, which matters if your downstream spec says 1280x720.

For a first pass, render the same prompt on wan-3.0 and gemini-omni-flash-1.1 at 720p. The cost is $2.00 for the pair, and that is a smaller bet than a seedance-2.5 run at $4.6224.

A batch example

Say you had 100 eight-second 720p clips a month on the old route. On wan-3.0 or gemini-omni-flash-1.1 that is 100 x $1.00 = $100.00 a month. On minimax-h3 at 768p it is 100 x $0.60 = $60.00, and on seedance-2.5 at 720p it is 100 x $4.6224 = $462.24. Sume reserves the estimated amount at submit and refunds a job that fails before generation, so a wrong first guess costs you the clips that completed and nothing else.

Run ten of each candidate before moving the whole workload. Ten clips on wan-3.0 is $10.00, and ten on gemini-omni-flash-1.1 is another $10.00, which is a cheap way to find out whether the duration cap, the aspect ratio set or the audio behavior breaks your pipeline.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume