Grok Imagine extension: duration is added seconds; Sume's chain
xAI's /v1/videos/extensions duration counts only the new seconds, so 10s plus 5 returns 15s. Sume has no extend call here: chain clips from the last frame.

On xAI's Grok Imagine video extension, duration controls the length of the extended portion only. xAI's own example: a 10-second input with duration set to 5 returns a 15-second video. Sume does not ship an extend endpoint in these docs, so the equivalent is a chain: pull the last frame of your clip with video frames, then generate the next clip with that frame as the first frame.
The xAI facts come from its Video Extension page, read 2026-10-02. The Sume side comes from Video frames and Video generation.
What does xAI's extension endpoint take?
The endpoint is POST /v1/videos/extensions on the model grok-imagine-video. You send a prompt that describes what happens next, a duration for the new footage, and a video as a public URL, a base64 data URI, or a Files API file_id. You poll /v1/videos/{request_id} for the extended video URL.
The page states no limit on source length, resolution, or file size, and no price or chaining cap, so none are repeated here.
| Question | xAI Grok Imagine | Sume |
|---|---|---|
| Dedicated extend call | Yes, /v1/videos/extensions | No |
| What duration means | Length of the added portion only | Length of the new clip you generate |
| Result | One video, original plus extension | A new clip you join yourself |
| Continuity tool | Source video passed in | First frame from the previous clip |
How do I continue a clip on Sume?
Two calls. First, extract a still near the end of your clip. Video frames takes one clip hosted on media.sume.com plus at[] times in seconds, and each time must be at least 0 and strictly less than the clip's duration, or the worker fails with frame_time_out_of_range. So for a 10-second clip, ask for 9.5, not 10. The extract is unbilled.
Second, generate the next clip with that image as the first_frame through POST /v1/videos. Only models whose supported_frame_images lists first_frame accept it; the catalog entry for seedance-2 does.
curl -X POST https://api.sume.com/v1/video-frames \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: last-frame-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/clip.mp4",
"at": [9.5]
}'What does the second call look like?
Read the frame URL from frames[0].url once resource_status is ready, then submit the next clip. Your prompt should describe the continuation, because the model only sees the frame, not the motion that led to it. Expect a visible seam in motion and sound, which is the main honest difference from a native extend.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "seedance-2",
"prompt": "The camera keeps pulling back as the runner crosses the finish line",
"frame_images": [{
"type": "image_url",
"image_url": {"url": "https://media.sume.com/artifacts/artf_demo/last-frame.jpg"},
"frame_type": "first_frame"
}]
}'What about the final video?
xAI hands you one longer file. On Sume you get separate clips, each its own job. Join them with Timeline 1.0 when you want one MP4. Because each clip is billed as its own generation, a three-link chain is three charges, plus the free frame extract.
If you were relying on xAI's extension to keep the same voice and soundtrack, test that on Sume before you commit: a new clip generated from a still has no memory of the previous audio.
When is the chain not good enough?
A chain restarts from a single image, so anything that lived only in the motion is lost: the speed of a pan, the direction of a walk, the audio bed. xAI's endpoint takes the whole source video as input, which at least gives the model the motion to continue. On Sume, describe the motion in the prompt and keep each link short.
For a long sequence, plan the shots rather than the seconds: three clips with different framings joined in a timeline read as an edit, while three clips that try to pretend to be one take expose every seam. Check the Video Router catalog for each model's duration limits before you pick the clip length, since limits are per model.
- xAI: one call, duration is the added seconds, you pass the whole video.
- Sume: frame extract (unbilled), then one generation per link, then a timeline join.
- Either way, store each job id before waiting.
Sources
Related posts
More in Developers
- MiniMax H3 Max lip sync audio_url: Sume media host only
Sume's MiniMax H3 Max Lip Sync takes a public HTTPS audio_url on the Sume media host, 5 to 14.8 seconds, max 10 MB. Other hosts are rejected.
- MiniMax H3 Max lip sync: 1080p max, no 2K, speed_tier ignored
Sume's MiniMax H3 Max Lip Sync offers 480p, 768p (default) and 1080p; 2K is not offered, and speed_tier is accepted but ignored.
- MiniMax H3 Max lip sync image aspect ratio: 0.4 to 2.5
The still for Sume's MiniMax H3 Max Lip Sync must have an aspect ratio between 0.4 and 2.5. That covers 9:16 through 21:9 if ratio means width over height.
- H3 Max Recast in Python: submit, poll and read the result
A Python script for Sume's h3-max-recast: submit with an Idempotency-Key, poll until the job is terminal, read the result. Cost per clip included.
Written by Sume