Do reference images or audio change Seedance 2.5's price?

No. On Sume only a reference video changes a Seedance 2.5 estimate. Reference images, reference audio and generate_audio add nothing. What fal says, and a test.

5 min readSume
All posts

No. On Sume a Seedance 2.5 request costs the same with 1 reference image or 9, with or without a reference audio clip, and with generate_audio on or off. The estimate has three inputs that matter: resolution, aspect ratio and output seconds. A reference video is the only reference type that changes it.

That matches what fal publishes for the model. Its text-to-video page says the audio flag "does not change your token count" because audio is generated in the same pass, and its reference-to-video page lists only input video duration in the token formula. Both pages were read on 2026-10-02.

What does the Seedance 2.5 price depend on?

Sume's video token estimate takes resolution, aspect ratio, duration and whether any reference video is present. It has no field for the number of images, for audio references, or for audio on or off. The table lists each input and what it does to the number.

Inputs that move a Seedance 2.5 estimate on Sume, read 2026-10-02
InputChanges the estimate?Effect
resolution (480p, 720p, 1080p)YesPixel count per frame; 1080p also uses a higher rate
aspect_ratioYesDifferent pixel dimensions per ratio
duration (4 to 30 s)YesLinear in output seconds
reference video(s)YesAdds 15 assumed seconds, then x0.6
reference imagesNoCounted against the request cap only
reference audioNoCounted against the request cap only
generate_audio true or falseNoAudio is produced in the same pass

What does fal say about audio and references?

On the text-to-video page, generate_audio is a boolean that is on by default, and fal states it does not change the token count. At 720p the page gives roughly $0.4730 per second, and its worked examples are about $2.31 for 5 seconds and $13.87 for 30 seconds.

On the reference-to-video page the formula adds the input video duration to the output duration and multiplies by 0.6 when video references are present, with up to 50 multimodal inputs per request. Images and audio appear in the input list but not in the formula. Sume prices from the same model family at list times 1.25, which is why its 720p second comes out near $0.59 instead of $0.47.

Does sending more reference images ever cost more?

Not in the estimate. It can still cost you in another way: Sume caps the combined count of reference images, videos and audio on a Seedance request at 12, and a request over the cap is rejected with an unsupported_capability error. The reference limits post covers the 12-input cap in detail.

Reference images also do not make a clip longer or sharper. They steer the content. If a run keeps drifting, adding a fourth identical reference is less useful than a clearer prompt tag, covered in the @image tag post.

What about Seedance 2.0 and the cheaper ids?

The same shape holds across the four Seedance ids Sume lists, because they share one pricing function and differ only in rate per 1,000 tokens: seedance-2.5 $0.0214, seedance-2 $0.014, seedance-2-fast $0.0112 and seedance-2-mini $0.007 at list, all multiplied by 1.25. For the 2.0 side of this question, see Does generate_audio change Seedance 2.0's price?.

If you want the cheapest sound-on test of a reference setup, use seedance-2-mini at 480p or 720p, then move the same request body to seedance-2.5 once the composition is right. The body does not need to change except for model.

How do you confirm it on your own account?

Submit the same prompt twice with generate_audio set to true and then false, and compare the usage.cost on the completed job responses. Per the Video generation docs, usage.cost is on the poll response. Expect identical numbers. Then add a reference video to one request and the number should jump by the 15-second rule in the crossover post.

One limit of this test: the Sume estimate is a reserve built from list prices, and the docs say the catalog row's billable_formula is the per-model truth. If a rate changes, the catalog changes first. Treat the table above as a snapshot dated 2026-10-02, and re-read the catalog row before you build a budget on it. A 30-second 720p clip with nine reference images and a voice clip should still show the same figure as the plain text-to-video request, about $17.34.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume