Grok Imagine mid-video frame pin vs Sume's first and last frames
xAI lets Grok Imagine 1.5 pin first, last and mid frames. Sume's Grok row takes one start still; this maps pinning to Sume rows with first and end frames.

Sume has no mid-video frame pin on any row, and its Grok Imagine row takes only a single start image. xAI's video guide says grok-imagine-video-1.5 can pin first, last and mid-video frames (xAI video guide, read 2026-10-10), so a script that relies on a mid frame needs a different structure on Sume.
The closest Sume equivalent is a row that accepts a first frame and an end frame, or a chain of two clips whose boundary frame you control.
What the two catalogs offer
On Sume the catalog describes frame control with two capability flags, image_to_video and end_frame. The Grok row sets end_frame to false and rejects end_image_url and last_frame_url. Rows that accept a start and end image are Seedance 2.5 and the 2.0 family, Kling Video v3 Pro, Wan 3.0, MiniMax H3, MiniMax H3 Max and Gemini Omni Flash 1.1 (Video Generation).
| Control | xAI Grok Imagine 1.5 | Sume Grok row | Sume rows with end frame |
|---|---|---|---|
| First frame | Yes | Yes (image_url) | Yes |
| Last frame | Yes | No | Yes |
| Mid-video frame | Yes | No | No |
| Reference images | Up to 14 | One | Seedance, Wan, MiniMax, Omni (not Kling) |
Two ways to approximate a mid frame
If the shot truly needs a landmark in the middle, split it:
- Make clip A from your first frame to the mid frame as its end frame, then clip B from the mid frame to the last frame. Join them on a timeline. Pick a row that accepts both frames, for example Kling Video v3 Pro (4 to 15 s) or Seedance 2.5 (4 to 30 s).
- Use reference images on a row that takes them (Wan 3.0 up to 10 images) and describe the order in the prompt. This is guidance, not a pin: Sume's docs say references are visual guidance, not accurate frames.
What changes in the request
On /v1/videos, frame_images carries a frame_type of first_frame or last_frame, and if you also send input_references, frame_images controls the mode and Sume processes the request as image-to-video. On the older Video Router, the same idea is image_url plus end_image_url.
Check supported_frame_images on the row first. For the Grok row, it lists only a first frame.
Cost and limits of the split approach
Two 6-second clips on Kling Video v3 Pro cost 12 seconds at the list rate of $0.112 per second without audio, billed at 1.25 times, so $0.14 per second: 12 x $0.14 = $1.68. The same 12 seconds on the Grok row at $0.0125 per second is $0.15, but it cannot take the second still as an end frame. The cheaper row does not do the job; pick by capability first, price second.
Sources
Related posts
More in Comparisons
- YouTube's four disclosure triggers checked against Sume's tools
Four YouTube disclosure triggers: real person, altered real footage, invented realistic scene, music. Each is checked against Sume generation and media tools.
- YouTube's resolution ladder: which rows Sume Timeline can render
YouTube lists eight 16:9 sizes from 426x240 to 7680x4320. Sume Timeline renders even edges from 256 to 2160, so 1080p, 720p, 480p and 360p fit; the rest do not.
- YouTube: upscale and repair need no AI label, realistic edits do
YouTube's page exempts sharpening, upscaling or repair, beauty filters and colour changes from AI disclosure. Realistic synthetic people or places need it.
- Sume vs Argil: AI avatar video and video agents compared
Argil makes AI-avatar and story videos with a chat agent, Director; Sume is a video agent with a multi-model API. Avatars, API, pricing, and limits compared.
Written by Sume