Design an episode's opening frame with Grok Imagine, then animate it

Make the first frame with x-ai/grok-image (9:16, about $0.025 a still on Sume), approve it, then pass it as first_frame to a video model. Costs and code.

5 min readSume
All posts

Grok Imagine on Sume is the image id x-ai/grok-image, and the cheapest way to use it for an episodic series is to make one approved still per episode and hand it to a video model as the first frame. A still is about $0.025 on Sume (read 2026-10-03), so six candidate openings for one episode cost $0.15 before any video is paid for.

xAI's own guide for grok-imagine-image-2.0 says generation can return up to 10 images per request, edits accept up to 5 source images as a public URL or base64 data URI, and image generation uses flat per-image pricing (read 2026-10-03). xAI's models page lists the model at $0.04 per image (read 2026-10-03). Sume's catalog id is a different surface with its own price line, and the repo does not say which Grok Imagine version sits behind it, so read the live entry rather than assuming parity.

What each side lets you set

The point of the comparison is not to pick a winner. It is to know which knobs exist where, so a recipe written against one does not silently break on the other.

Grok Imagine image options, xAI and Sume (read 2026-10-03)
ItemxAI grok-imagine-image-2.0Sume x-ai/grok-image
Images per requestUp to 10Check n in GET /v1/images/models
Edit sourcesUp to 5, URL or base64Check input_references in the same call
Price per image$0.04 (xAI models page)$0.025 billable in the Sume rate card
Aspect ratiosNot on the pages read13, including 9:16 and 16:9; no 4:5
OutputNot on the pages readHosted URL in data[].url

The opening-frame workflow

Start with the shot, not the style. An episode opening has three jobs: show where we are, show who is in frame, and leave room for the first movement. Write the prompt for that single frame in 9:16, and keep the subject away from the edges, because the video model will move the camera or the subject out of that frame.

Generate a few candidates, pick one, and only then pay for motion. A still is cheap and a video clip is not, so every bad composition rejected at the still stage saves a video job. The hand-off to video uses the documented frame_images field with a first_frame entry; read supported_frame_images on GET /v1/videos/models to confirm the video model you chose accepts one.

import os, requests

API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

img = requests.post(f"{API}/v1/images", headers=H, json={
    "model": "x-ai/grok-image",
    "prompt": "Rain-slick alley at dusk, neon sign flickering, a courier in a "
              "yellow coat standing left of centre, empty space to the right",
    "aspect_ratio": "9:16",
})
img.raise_for_status()
still = img.json()["data"][0]["url"]
print("approve this still:", still)

if input("animate it? [y/N] ").strip().lower() == "y":
    vid = requests.post(f"{API}/v1/videos",
        headers={**H, "Idempotency-Key": "s1e1-opening"},
        json={"model": "seedance-2.5", "duration": 6, "aspect_ratio": "9:16",
              "prompt": "The courier looks up as the sign flickers; slow push-in.",
              "frame_images": [{"type": "image_url", "image_url": {"url": still},
                                "frame_type": "first_frame"}]})
    print(vid.status_code, vid.json())

What to do for a season

Write one prompt template for the series with the fixed parts (look, palette, wardrobe, lens language) and a short per-episode slot for the scene. Reuse the template for every episode's opening so the stills belong together, then store the approved still next to the episode record.

Keep the aspect ratio in the supported list. Sume's Grok ratios are 2:1, 20:9, 19.5:9, 16:9, 4:3, 3:2, 1:1, 2:3, 3:4, 9:16, 9:19.5, 9:20 and 1:2, so a 4:5 feed crop is not in it. Make the still at 9:16 and crop later, or choose a model whose list includes 4:5.

Run a small acceptance test on episode one before you scale. Generate the still, animate it for the shortest duration the model accepts, and compare frame zero of the clip with the still. If the first frame matches but the subject drifts by second three, shorten the shots or move the camera instead of the subject. If frame zero itself differs, the model is treating your image as a loose reference, and you should pick another model from the catalog rather than rewriting the prompt.

Budget it. Six candidate stills per episode over eight episodes is 48 stills, which is $1.20 at $0.025 each. The video clips under them are the real cost, and they vary by model and duration, so price them from the live catalog before you commit. A habit that helps: record the still's URL and the prompt in the same row, so a retake of the opening can start from the approved composition instead of from scratch.

What Sume does and does not do

Sume lists x-ai/grok-image on the Image API and returns a hosted URL you can pass straight to a video request. It does not promise the video will keep the still's details, and it applies the reference-image count its own catalog descriptor lists, not xAI's five. Treat the xAI numbers above as context for the vendor's own API, not as Sume behaviour.

Sources

Related posts

More in Models

All Models posts

Written by Sume