grok-imagine-video-1.5 on Sume: needs an image, seven fields refused
grok-imagine-video-1.5 is image-to-video only on Sume: send one image, no end frame, no reference video or audio, no aspect_ratio, no generate_audio.

On Sume, grok-imagine-video-1.5 is an image-to-video model. A request without image_url, first_frame_url or one reference_image_urls entry is refused with "model grok-imagine-video-1.5 requires image_url, first_frame_url, or one reference_image_urls entry." Seven other fields are also refused by name, so a text-only or multi-reference prompt must go to a different model.
The envelope
The catalog lists 480p and 720p, durations from 4 to 15 seconds, no end frame, no reference videos or audios and no audio output. The validation then goes further than the capability flags: a second reference image is refused with a message that says to use image_url for the source frame, and the fields in the table below each return "X is not supported by model grok-imagine-video-1.5."
| Field | Result |
|---|---|
| end_image_url | Refused |
| last_frame_url | Refused |
| reference_video_urls | Refused |
| reference_audio_urls | Refused |
| bitrate_mode | Refused |
| aspect_ratio | Refused |
| generate_audio | Refused: no audio field |
| Second reference_image_urls entry | Refused: use image_url for the source frame |
Shape the frame with the image
Because aspect_ratio is refused for this model, shape the frame with the image itself: crop or pad the picture to the ratio you want before you send it. video_frames extracts a still from a stored clip, which is a handy source for a continuation.
Minimal valid request
The body has a model, a prompt, an image and a resolution. It prints the request and only sends when you set SEND=1, because a valid call is billed.
import os
import requests
body = {
"model": "grok-imagine-video-1.5",
"prompt": "The camera drifts right as steam rises from the cup",
"image_url": "https://example.com/cup.png",
"resolution": "720p",
"duration": 6,
}
print(body)
if os.environ.get("SEND") == "1":
r = requests.post("https://api.sume.com/v1/video-router/generate",
headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"]},
json=body, timeout=60)
print(r.status_code)
When to pick something else
If you need an ending frame, use kling-3, a Seedance model or Wan 3.0, which have end-frame support. If you need references or audio, use a model whose capabilities list them. If you want a text-only prompt, pick any text-to-video model. Grok's role in this catalog is narrow: animate one image, at 480p or 720p.
A few habits avoid the refusals. Keep one request builder per model family rather than one generic builder with every field. Read the model's capability flags from the catalog endpoint before you add an optional field. And log the 400 message verbatim: for this model each refusal names the field, so the fix is usually a single deleted line.
Cost matters as well. The model is a 4 to 15 second image-to-video option at two resolutions, so check the per-second price in the catalog entry against Seedance or Wan before you commit a batch to it; the right choice depends on the look you need, not on the cheapest row.
Finally, because there is no end frame and no reference list, continuity across shots has to come from the images you feed in. Extract the last still of the previous clip and use it as the next image_url, and keep the prompt focused on motion rather than re-describing what the image already shows.
Sources
Related posts
More in Models
- H3 Max 1080p regenerates from 768p: the price of the extra step
MiniMax says its 2K path regenerates in context. Sume documents H3 Max 1080p as a latent refinement of 768p, $0.20 vs $0.10 per second. When the doubling pays.
- H3 or H3 Max? Pick the Sume row by job, with per-clip prices
Both read references and make stereo audio. H3 is $0.075 per second at 768p, H3 Max $0.10 with a 1080p row. A job-by-job guide with 10-second prices.
- higgsfield-genjutsu missing from /v1/videos/models: why
higgsfield-genjutsu is in the Sume video catalog only when its provider is configured. It is Motion Transfer: one video_url plus 1-8 images, 480p or 720p.
- Choose a Sume video model in five questions
Five questions pick a Sume video model: length, resolution, source media, audio toggle and price. One dated table maps each answer to model ids.
Written by Sume