Seedance vs Grok Imagine: inputs, length, sound and cost
Grok Imagine Video 1.5 animates one image into a silent clip, billed per second. Seedance 2.x adds text, end frames, references and sound, billed by token.

Seedance and Grok Imagine differ most in how much a request can steer and how the meter runs. Seedance 2.0 and 2.5 accept reference images, videos and audio, generate sound by default, reach 1080p, and seedance-2.5 runs to 30 seconds; they are billed by video tokens, so cost scales with pixels as well as seconds. The Grok model takes none of those inputs and bills a flat $0.0125 per second.
This compares the models as Sume's POST /v1/videos serves them, not the consumer apps, from Sume's video catalog and pricing code and the Sume API reference, read 2026-09-28.
Where do Seedance and Grok Imagine differ?
The ids are seedance-2 (Seedance 2.0), seedance-2.5 and grok-imagine-video-1.5. grok-imagine-video-1.5 only animates one still, with no text-only mode; the single-model posts for it and for Seedance 2.5 carry the full field lists. The rows below are the ones that separate them:
| `seedance-2` | `seedance-2.5` | `grok-imagine-video-1.5` | |
|---|---|---|---|
| Reference images, videos, audio | Yes | Yes | No |
| Generated sound | On by default | On by default | Never |
| Longest clip | 15 s | 30 s | 15 s |
| Top resolution | 1080p | 1080p | 720p |
| Meter | $0.0175 per 1,000 video tokens | $0.02675 per 1,000 video tokens (more at 1080p) | $0.0125 per second |
What can Seedance steer that Grok Imagine can't?
Seedance takes reference files next to the prompt: images for a character or product look, a video for motion, audio for a voice or beat. Audio and video references are honored by the Seedance 2.x models. The single-still model has no field for any of them, so its look comes from that still and the prompt alone.
Audio follows the same split. Seedance writes dialogue, music or effects into the clip unless generate_audio is set to false; the other returns no audio track, so a voiceover or soundtrack is a separate step.
How does token billing compare with per-second billing?
The per-second meter ignores frame size. Seedance bills video tokens: width × height × seconds × 24 fps ÷ 1,024, which is 21,600 tokens per second of 720p 16:9 video, so a larger frame costs more even at the same length. In current code, Sume's Seedance estimate has no audio term, so sound on or off does not change it. Seedance video token pricing works the formula at other sizes.
| Clip | `grok-imagine-video-1.5` | `seedance-2` | `seedance-2.5` |
|---|---|---|---|
| 5 s | $0.07 | $1.89 | $2.89 |
| 10 s | $0.13 | $3.78 | $5.78 |
| 15 s | $0.19 | $5.67 | $8.67 |
Which one should I use?
- Use
grok-imagine-video-1.5when one still plus silent motion is enough and volume makes the per-second meter matter: 10 seconds at 720p reserves $0.13. - Use Seedance when references must hold a look, the clip needs its own sound, runs past 15 seconds (
seedance-2.5) or needs 1080p. The same 10 seconds at 720p reserves $3.78 onseedance-2. - For Seedance at lower token rates, see Seedance 2.0 Fast vs Mini; for the single-still model against Kling, see Kling vs Grok Imagine.
Sources
Related posts
More in Models
- Seedance vs Seedream: ByteDance's video and image models
Seedance is ByteDance Seed's video model family; Seedream is its image model family. What each makes, takes, and costs through one API.
- Seedance vs Sora 2: specs side by side after the shutdown
Sora 2 shut down on September 24, 2026. Its clip lengths, frame sizes, inputs, extensions and edits beside Seedance 2.0 and 2.5, which you can still call.
- Sora MCP server: why it stopped working and what to use
A Sora MCP server can't make video now: OpenAI shut the Sora API down on September 24, 2026. Point your MCP client at other video models instead.
- Text to speech in Japanese: set the language to ja
For Japanese text to speech, send the script in kana and kanji and set the language to ja. Left out, Sume can read a kanji-only line as English.
Written by Sume