Text or image start: Seedance 2.5 is $5.78 either way (10 s, 720p)
Starting a Sume video from an image does not change its price (model, resolution, ratio, seconds). Which models need an image, plus one H3 exception.

No. On Sume, a text-to-video start and an image-to-video start are the same price for the same model, resolution, aspect ratio and length. A 10-second 720p 9:16 clip on Seedance 2.5 is $5.78 whichever way you start it. The grid that prices a clip has four inputs: model, resolution, aspect ratio (or audio setting for Kling), and seconds. There is no start-mode input. The exceptions are models that need an image, and one reference-image overage on MiniMax H3.
What each model accepts as a start
The mode comes from the fields you send. frame_images with first_frame or last_frame gives image-to-video. input_references gives reference-to-video, where the images guide style or content but are not exact frames. If you send both, frame_images decides and Sume treats the request as image-to-video.
| Model | Text start | Image start | First plus last frame |
|---|---|---|---|
| seedance-2.5, 2, Fast, Mini | yes | yes | yes |
| wan-3.0 | yes | yes | yes |
| kling-3 | yes | yes | yes |
| minimax-h3, h3-max | yes | yes | yes |
| gemini-omni-flash-1.1 | yes | yes | yes |
| higgsfield-genjutsu | no, source video required | no | no |
| h3-max-recast | no, source video required | no | no |
Where the price does move
- MiniMax H3 reference-to-video: the first five reference images are free and each extra image is $0.08 list, so $0.10 billable at 1.25.
- MiniMax H3 Max reference-to-video is billed as output seconds at list x 1.25; the catalog states that fal reference-token overage is not reserved.
- Kling 3: the audio setting changes the price ($1.40 silent against $2.10 with sound for 10 s at 720p), but the start mode does not.
- Genjutsu and Recast are source-video tools and have their own per-second rows.
Why the start still matters for cost
Equal price per clip does not make the two starts equal in practice. An image start constrains the opening frame, which is what you pay to control. If a first frame is wrong, you repay the whole clip, so a cheap still-image check before the video job is the cheaper place to catch mistakes. A text start leaves the model free to choose the first frame, and you find out only after the full job. For Seedance 2.5 at 720p, a 10 s retry is $5.78, while a 480p draft at 10 s is $2.69.
A rule of thumb
Use a text start for exploration and an image start when a fixed product, face or layout must open the clip. Pin the model in Video generation, read supported_frame_images from GET /v1/videos/models, and send frame_images for exact frames. Use input_references only when you want guidance rather than an exact frame.
Check it before you run it
Every figure above is a catalog list price times 1.25, rounded up to the cent, as of 2026-10-08. Catalogs change, so before a large batch, read the current model entry in the docs and recompute the one line that matters for your case. Write the arithmetic next to the job in your own notes: list rate, seconds or characters, multiplier, rounding. If the result differs from the wallet charge by more than a cent, the catalog entry has changed, and the docs page is the place to find out why.
Run one small job first. Submit a single request with the pinned model id and the settings in the tables, poll the returned polling_url until it finishes, and compare the charge with your estimate. Then scale up. Using a pinned id for the test matters, because sume/auto never names the family, so you cannot tie its charge to the row you priced.
Sources
Related posts
More in Comparisons
- TikTok Spark vs non-Spark ad length: no restriction vs 10 minutes
TikTok's spec page lists no duration restriction for Spark Ads and up to 10 minutes for non-Spark. Format, size and caption differences, and Sume output.
- TTS per 1,000 characters: Sume $0.0475 vs MAI-Voice-2.1 $0.022
Sume TTS lists $0.0475 per 1,000 characters; Microsoft lists MAI-Voice-2.1 at $22 per 1M and Flash at $15 per 1M. List prices side by side, with the job math.
- VEED green screen API per 30 frames: Sume has no video matting
VEED prices Green Screen at $0.025 per 30 frames. Sume does not remove video backgrounds; its filter route only dims, crops and applies ffmpeg filters.
- VEED Lip Sync 2.0 at $0.07 a second vs Sume still plus audio
VEED Lip Sync 2.0 is $0.07 a second and works video to video. Sume's lip sync starts from a still and 5 to 14.8 seconds of audio. They solve different jobs.
Written by Sume