Design an episode's opening frame with Grok Imagine, then animate it
Make the first frame with x-ai/grok-image (9:16, about $0.025 a still on Sume), approve it, then pass it as first_frame to a video model. Costs and code.

Grok Imagine on Sume is the image id x-ai/grok-image, and the cheapest way to use it for an episodic series is to make one approved still per episode and hand it to a video model as the first frame. A still is about $0.025 on Sume (read 2026-10-03), so six candidate openings for one episode cost $0.15 before any video is paid for.
xAI's own guide for grok-imagine-image-2.0 says generation can return up to 10 images per request, edits accept up to 5 source images as a public URL or base64 data URI, and image generation uses flat per-image pricing (read 2026-10-03). xAI's models page lists the model at $0.04 per image (read 2026-10-03). Sume's catalog id is a different surface with its own price line, and the repo does not say which Grok Imagine version sits behind it, so read the live entry rather than assuming parity.
What each side lets you set
The point of the comparison is not to pick a winner. It is to know which knobs exist where, so a recipe written against one does not silently break on the other.
| Item | xAI grok-imagine-image-2.0 | Sume x-ai/grok-image |
|---|---|---|
| Images per request | Up to 10 | Check n in GET /v1/images/models |
| Edit sources | Up to 5, URL or base64 | Check input_references in the same call |
| Price per image | $0.04 (xAI models page) | $0.025 billable in the Sume rate card |
| Aspect ratios | Not on the pages read | 13, including 9:16 and 16:9; no 4:5 |
| Output | Not on the pages read | Hosted URL in data[].url |
The opening-frame workflow
Start with the shot, not the style. An episode opening has three jobs: show where we are, show who is in frame, and leave room for the first movement. Write the prompt for that single frame in 9:16, and keep the subject away from the edges, because the video model will move the camera or the subject out of that frame.
Generate a few candidates, pick one, and only then pay for motion. A still is cheap and a video clip is not, so every bad composition rejected at the still stage saves a video job. The hand-off to video uses the documented frame_images field with a first_frame entry; read supported_frame_images on GET /v1/videos/models to confirm the video model you chose accepts one.
import os, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
img = requests.post(f"{API}/v1/images", headers=H, json={
"model": "x-ai/grok-image",
"prompt": "Rain-slick alley at dusk, neon sign flickering, a courier in a "
"yellow coat standing left of centre, empty space to the right",
"aspect_ratio": "9:16",
})
img.raise_for_status()
still = img.json()["data"][0]["url"]
print("approve this still:", still)
if input("animate it? [y/N] ").strip().lower() == "y":
vid = requests.post(f"{API}/v1/videos",
headers={**H, "Idempotency-Key": "s1e1-opening"},
json={"model": "seedance-2.5", "duration": 6, "aspect_ratio": "9:16",
"prompt": "The courier looks up as the sign flickers; slow push-in.",
"frame_images": [{"type": "image_url", "image_url": {"url": still},
"frame_type": "first_frame"}]})
print(vid.status_code, vid.json())What to do for a season
Write one prompt template for the series with the fixed parts (look, palette, wardrobe, lens language) and a short per-episode slot for the scene. Reuse the template for every episode's opening so the stills belong together, then store the approved still next to the episode record.
Keep the aspect ratio in the supported list. Sume's Grok ratios are 2:1, 20:9, 19.5:9, 16:9, 4:3, 3:2, 1:1, 2:3, 3:4, 9:16, 9:19.5, 9:20 and 1:2, so a 4:5 feed crop is not in it. Make the still at 9:16 and crop later, or choose a model whose list includes 4:5.
Run a small acceptance test on episode one before you scale. Generate the still, animate it for the shortest duration the model accepts, and compare frame zero of the clip with the still. If the first frame matches but the subject drifts by second three, shorten the shots or move the camera instead of the subject. If frame zero itself differs, the model is treating your image as a loose reference, and you should pick another model from the catalog rather than rewriting the prompt.
Budget it. Six candidate stills per episode over eight episodes is 48 stills, which is $1.20 at $0.025 each. The video clips under them are the real cost, and they vary by model and duration, so price them from the live catalog before you commit. A habit that helps: record the still's URL and the prompt in the same row, so a retake of the opening can start from the approved composition instead of from scratch.
What Sume does and does not do
Sume lists x-ai/grok-image on the Image API and returns a hosted URL you can pass straight to a video request. It does not promise the video will keep the still's details, and it applies the reference-image count its own catalog descriptor lists, not xAI's five. Treat the xAI numbers above as context for the vendor's own API, not as Sume behaviour.
Sources
Related posts
More in Models
- Hebrew and Georgian text to speech API: he and ka on Sume TTS
Cartesia lists Hebrew (he) and Georgian (ka) on Sonic 3.6, 3.5 and 3. What to send to Sume TTS, why the library has no voice for them, and what a script costs.
- Higgsfield Soul 2 custom_reference_id is not a Sume parameter
Higgsfield's Soul 2 API takes custom_reference_id for a saved character. Sume's Soul is text-only; for a consistent character use a model that takes references.
- Soul 2 lookbook batch of 4: Higgsfield Soul on Sume
Generate a four-image lookbook set per request with Higgsfield Soul 2: batch_size 1 or 4, 720p or 1080p, seven ratios, text-to-image only.
- Higgsfield Soul on Sume: text to image, 1 or 4 images, 720p or 1080p
Soul is a text-to-image row in Sume's image catalog: seven aspect ratios, no references, 1 or 4 images per call, 720p or 1080p. Limits and price.
Written by Sume