Text or image start: Seedance 2.5 is $5.78 either way (10 s, 720p)

Starting a Sume video from an image does not change its price (model, resolution, ratio, seconds). Which models need an image, plus one H3 exception.

5 min readSume
All posts

No. On Sume, a text-to-video start and an image-to-video start are the same price for the same model, resolution, aspect ratio and length. A 10-second 720p 9:16 clip on Seedance 2.5 is $5.78 whichever way you start it. The grid that prices a clip has four inputs: model, resolution, aspect ratio (or audio setting for Kling), and seconds. There is no start-mode input. The exceptions are models that need an image, and one reference-image overage on MiniMax H3.

What each model accepts as a start

The mode comes from the fields you send. frame_images with first_frame or last_frame gives image-to-video. input_references gives reference-to-video, where the images guide style or content but are not exact frames. If you send both, frame_images decides and Sume treats the request as image-to-video.

Start modes in the Sume video catalog, read 2026-10-08
ModelText startImage startFirst plus last frame
seedance-2.5, 2, Fast, Miniyesyesyes
wan-3.0yesyesyes
kling-3yesyesyes
minimax-h3, h3-maxyesyesyes
gemini-omni-flash-1.1yesyesyes
higgsfield-genjutsuno, source video requirednono
h3-max-recastno, source video requirednono

Where the price does move

  • MiniMax H3 reference-to-video: the first five reference images are free and each extra image is $0.08 list, so $0.10 billable at 1.25.
  • MiniMax H3 Max reference-to-video is billed as output seconds at list x 1.25; the catalog states that fal reference-token overage is not reserved.
  • Kling 3: the audio setting changes the price ($1.40 silent against $2.10 with sound for 10 s at 720p), but the start mode does not.
  • Genjutsu and Recast are source-video tools and have their own per-second rows.

Why the start still matters for cost

Equal price per clip does not make the two starts equal in practice. An image start constrains the opening frame, which is what you pay to control. If a first frame is wrong, you repay the whole clip, so a cheap still-image check before the video job is the cheaper place to catch mistakes. A text start leaves the model free to choose the first frame, and you find out only after the full job. For Seedance 2.5 at 720p, a 10 s retry is $5.78, while a 480p draft at 10 s is $2.69.

A rule of thumb

Use a text start for exploration and an image start when a fixed product, face or layout must open the clip. Pin the model in Video generation, read supported_frame_images from GET /v1/videos/models, and send frame_images for exact frames. Use input_references only when you want guidance rather than an exact frame.

Check it before you run it

Every figure above is a catalog list price times 1.25, rounded up to the cent, as of 2026-10-08. Catalogs change, so before a large batch, read the current model entry in the docs and recompute the one line that matters for your case. Write the arithmetic next to the job in your own notes: list rate, seconds or characters, multiplier, rounding. If the result differs from the wallet charge by more than a cent, the catalog entry has changed, and the docs page is the place to find out why.

Run one small job first. Submit a single request with the pinned model id and the settings in the tables, poll the returned polling_url until it finishes, and compare the charge with your estimate. Then scale up. Using a pinned id for the test matters, because sume/auto never names the family, so you cannot tie its charge to the row you priced.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume