Grok Imagine video edit on xAI vs the video_url error on Sume
xAI edits video with grok-imagine-video; Sume accepts video_url only on Gemini Omni Flash 1.1. The exact error, the edit limits and what to send instead.

Sume cannot edit a video with a Grok Imagine model: video edit on Sume runs only on gemini-omni-flash-1.1, and any other model that receives a video_url gets a 400 validation error. xAI's own guide says editing an existing clip with a text prompt is done on the grok-imagine-video model, while grok-imagine-video-1.5 is the generation model (xAI video guide, read 2026-10-10).
That split matters in launch week, because Grok Imagine Video 1.5 Lite arrived on top of an existing lineup and people are porting scripts that mix both ids. A script that sends a source clip to a Grok id will fail on Sume before any money is reserved.
What xAI documents for editing
The xAI guide lists three modes for the Grok Imagine family: text-to-video, image-to-video (the still becomes the starting point) and video editing, where you modify an existing video with a text prompt on grok-imagine-video. It also states that grok-imagine-video-1.5 takes up to 14 reference images and up to 3 voice references, and that requests are asynchronous: you start a request, poll with the returned id, and fetch the finished video URL.
The guide does not give an input-length cap for the edit model, so this post makes no claim about one. For prices, xAI's models page shows $0.020 per second for grok-imagine-video-1.5-lite and $0.080 per second for grok-imagine-video-1.5 (read 2026-10-10); it does not give an edit price in the text we read.
What Sume does with a video_url
On Sume the Video Router schema checks the model before it reserves any balance. If the body carries video_url and the model is not gemini-omni-flash-1.1, validation fails with this message:
video_url (edit) is supported only by model gemini-omni-flash-1.1.
The Grok row in Sume's catalog is grok-imagine-video-1.5, and it is an image-to-video row only. It needs image_url, first_frame_url or one reference_image_urls entry, and it rejects end_image_url, last_frame_url, reference_video_urls, reference_audio_urls, bitrate_mode, aspect_ratio and generate_audio. So a Grok edit request fails twice over: wrong model for video_url, and no text-to-video input on that row at all.
| Input | xAI Grok Imagine | Sume row |
|---|---|---|
| Source video to edit | grok-imagine-video | Only gemini-omni-flash-1.1 (video_url) |
| Start image | grok-imagine-video-1.5 | grok-imagine-video-1.5: image_url or first_frame_url |
| Reference images | Up to 14 on 1.5 | One image on the Grok row |
| Voice references | Up to 3 on 1.5 | Not accepted on the Grok row |
| Text only | Yes | Not on the Grok row; it needs an image |
The Sume edit route that does exist
Gemini Omni Flash 1.1 edit takes a prompt plus video_url. Resolution is optional and defaults to 720p. Do not send aspect_ratio or duration: the output keeps the source framing, and a duration is at most a reserve-estimate hint. You also cannot combine video_url with image_url, end_image_url or reference_*_urls; the schema answers with Use either video_url (edit) or frame/reference_*_urls fields, not both.
Native audio is always on for that row, and the API rejects generate_audio: false. Sume bills the provider list times 1.25 per output second by resolution; the 720p list rate in the repo is $0.10 per second, so $0.125 billed (check /v1/videos/models for the live rate).
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: edit-demo-001" \
-d '{"model":"gemini-omni-flash-1.1","prompt":"Make the jacket red. Keep everything else the same.","video_url":"https://example.com/clip.mp4","resolution":"720p","mode":"async"}'Porting checklist
Before you move an xAI edit job, run through these checks:
- Swap the model id to
gemini-omni-flash-1.1and keep the instruction short and specific. - Remove
aspect_ratioanddurationfrom edit requests. - Keep the source clip reachable over public HTTPS, as Sume's errors page asks for input media.
- Send an
Idempotency-Keyso a retry returns the original job.
Sources
Related posts
More in Comparisons
- Hedra batch_size 1-8 variations vs Sume's one job per request
Hedra's API returns 1 to 8 variations per request; Sume avatar videos are one job each. How to get variations on Sume with previews and keys (read 2026-10-10).
- Hedra Character 3 vs Sume Avatar 1.0: per-second cost, 30 and 60 s
Hedra lists Character 3 at 2.5 to 6.25 cents per second; Sume Avatar 1.0 is $0.184 to $0.55 per second at 720p. Totals at 30 and 60 s (read 2026-10-10).
- HeyGen drops avatar_look_id on Oct 31: what to store on Sume
HeyGen's changelog deprecates avatar_look_id, character_id and character_type on Oct 31, 2026. See which fields a Sume Avatar 1.0 integration keeps instead.
- HeyGen avatar_video.fail has no fields: what Sume sends on failure
HeyGen says to treat avatar_video.fail as a signal and re-read the video. Sume's job.failed carries status ERROR and an error object. Read 2026-10-10.
Written by Sume