Sora 2 vs Grok Imagine video 1.5: text-first or image-first
Sora took a text prompt. grok-imagine-video-1.5 on Sume requires an image_url, runs 4 to 15 seconds at 480p or 720p, and bills $0.0125 per second.

Grok Imagine video 1.5 is not a drop-in text-to-video replacement for Sora: on Sume it is image-to-video only, so every request needs an image_url. If your Sora calls started from a text prompt alone, generate or pick a still first, then animate it for 4 to 15 seconds at 480p or 720p.
Why this is a different job
OpenAI's guide, read 2026-10-07, says the Sora 2 models and Videos API were shut down on September 24, 2026. Its deprecations page lists no replacement, so the choice of successor is yours.
A practical way to find out which camp your traffic is in is to count how many of your saved Sora requests included an input image. Those are candidates for grok-imagine-video-1.5 as they stand. The rest started from words alone and need a text-capable id or a still generated beforehand.
The two side by side
| Property | Sora 2 (OpenAI guide) | grok-imagine-video-1.5 on Sume |
|---|---|---|
| Starting point | Text prompt | image_url is required |
| Length | 16 or 20 second generations | 4 to 15 seconds |
| Resolution | 1080p on sora-2-pro only | 480p or 720p |
| Audio | Not covered here | No audio |
| Aspect ratio and end frame | Not compared here | Neither field is accepted |
| Billed price per second | API ended | $0.0125 |
What it costs
At $0.0125 per second billed (list price times 1.25), a 10 second clip comes to $0.125, which Sume rounds up to $0.13 on the reservation. That is the lowest per-second number among the video ids Sume lists, which makes it a sensible cheap animatic step before an expensive final pass.
When it fits your Sora prompts
Because the model animates the still you give it, the frame fixes the framing and the aspect. That suits product shots, portraits and illustrations; it does not suit prompts that rely on a model inventing the first frame.
Two-step flows are common once a model is image-first. First generate the still with an image model in the same account, check it, then animate it. You pay for the still once and can animate it several times at $0.0125 per second, trying different prompts on the same frame until the motion reads.
- Product still to 5 second loop: a good fit.
- Character with a spoken line: not a fit, there is no audio.
- Vertical 9:16 output: send a 9:16 still, since there is no aspect_ratio field.
If you need text-only
For text-first work, compare the other seven ids and pick one that takes a prompt alone. The request shape is the same POST to /v1/videos; only the model id and the required image_url differ. See the video docs for the field rules.
Whichever route you take, keep the model id in configuration. The shutdown showed that a vendor can retire an API with six months of notice, and an id you can change in one place is the cheapest insurance against the next one.
Sources
Related posts
More in Comparisons
- Sora 2 vs Wan 3.0 after the API shutdown: length, size and price
Sora 2 and its Videos API ended 2026-09-24. Wan 3.0 on Sume does 2 to 30 seconds at 480p to 1080p; limits and billed price per second, side by side.
- Pilot STT in 3 languages: language_code hint vs auto-detect, 30 clips
MAI-Transcribe-2-Streaming advertises 60 languages; here is a 30-clip pilot on Sume STT that compares a language_code hint with auto-detect for about 60 cents.
- Sume avatar product surcharge: 1, 1.3 or 3 cents a second by tier
A product on a Sume avatar clip adds 1 cent a second on Standard, 1.3 cents on Plus and 3 cents on Max. On 60 seconds: $0.60, $0.78 or $1.80.
- Sume Free vs Pro: what $40 a month adds beyond three more job slots
Free is $0 with one concurrent job and image models only. Pro is $40 a month for four slots, 100+ video models, API, CLI and MCP access, and full Sume Agent.
Written by Sume