Sora 2 vs Grok Imagine video 1.5: text-first or image-first

Sora took a text prompt. grok-imagine-video-1.5 on Sume requires an image_url, runs 4 to 15 seconds at 480p or 720p, and bills $0.0125 per second.

4 min readSume
All posts

Grok Imagine video 1.5 is not a drop-in text-to-video replacement for Sora: on Sume it is image-to-video only, so every request needs an image_url. If your Sora calls started from a text prompt alone, generate or pick a still first, then animate it for 4 to 15 seconds at 480p or 720p.

Why this is a different job

OpenAI's guide, read 2026-10-07, says the Sora 2 models and Videos API were shut down on September 24, 2026. Its deprecations page lists no replacement, so the choice of successor is yours.

A practical way to find out which camp your traffic is in is to count how many of your saved Sora requests included an input image. Those are candidates for grok-imagine-video-1.5 as they stand. The rest started from words alone and need a text-capable id or a still generated beforehand.

The two side by side

Sora 2 against grok-imagine-video-1.5 on Sume (read 2026-10-07)
PropertySora 2 (OpenAI guide)grok-imagine-video-1.5 on Sume
Starting pointText promptimage_url is required
Length16 or 20 second generations4 to 15 seconds
Resolution1080p on sora-2-pro only480p or 720p
AudioNot covered hereNo audio
Aspect ratio and end frameNot compared hereNeither field is accepted
Billed price per secondAPI ended$0.0125

What it costs

At $0.0125 per second billed (list price times 1.25), a 10 second clip comes to $0.125, which Sume rounds up to $0.13 on the reservation. That is the lowest per-second number among the video ids Sume lists, which makes it a sensible cheap animatic step before an expensive final pass.

When it fits your Sora prompts

Because the model animates the still you give it, the frame fixes the framing and the aspect. That suits product shots, portraits and illustrations; it does not suit prompts that rely on a model inventing the first frame.

Two-step flows are common once a model is image-first. First generate the still with an image model in the same account, check it, then animate it. You pay for the still once and can animate it several times at $0.0125 per second, trying different prompts on the same frame until the motion reads.

  • Product still to 5 second loop: a good fit.
  • Character with a spoken line: not a fit, there is no audio.
  • Vertical 9:16 output: send a 9:16 still, since there is no aspect_ratio field.

If you need text-only

For text-first work, compare the other seven ids and pick one that takes a prompt alone. The request shape is the same POST to /v1/videos; only the model id and the required image_url differ. See the video docs for the field rules.

Whichever route you take, keep the model id in configuration. The shutdown showed that a vendor can retire an API with six months of notice, and an id you can change in one place is the cheapest insurance against the next one.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume