Seedance 2.5 reference counts: Together's 30/10/10 vs Sume docs
Together lists Seedance 2.5 at 30 images, 10 videos and 10 audio clips per pass. Sume's docs name reference types for it but give no such counts.

Together AI's page for Seedance 2.5 says the model takes up to 30 images, 10 video clips and 10 audio clips in one pass, 50 in all, and generates 30-second audio-video clips. Sume's docs confirm seedance-2.5 accepts 4 to 30 seconds at 480p, 720p and 1080p, and that the Seedance 2.x models take audio and video references. The docs give no per-request counts for Seedance, so do not assume Together's numbers hold on Sume.
A quota on a vendor page is a statement about that vendor's endpoint. A router can pass fewer, or enforce its own caps.
What each source gives
| Limit | Together AI page | Sume docs |
|---|---|---|
| Clip length | 30 s per pass | 4 to 30 s |
| Image references | Up to 30 | Image type accepted; no count stated |
| Video references | Up to 10 | Video type accepted; no count stated |
| Audio references | Up to 10 | Audio type accepted; no count stated |
The one model where Sume does state counts
For gemini-omni-flash-1.1, the Video Router docs set reference caps: up to 10 images and up to 3 videos, each up to 3 seconds, with no audio references. Those numbers are for that model only. They show that a router's caps can be far lower than a headline vendor figure, which is why you should not carry any one model's counts to another.
A test that settles it
Build the request you need with the real number of references and send it to a test key. Start with a small clip at 480p to keep the test cheap. If it passes, check the output for every reference you meant to use, since accepting a reference is not the same as using it well.
Also read supported_input_references from the catalog at submit time rather than copying this table. The docs tell you to check the catalog because limits differ by model and change as models are added.
Why the counts differ in kind
Together's figures describe the model on Together's endpoint. A Sume request passes through the Video Router, which validates fields before it forwards them. The router can accept fewer references than the model can use. Neither number is wrong. They answer different questions, so the one that governs your cost and success is the one on the platform you call.
For a design that needs many references, such as a character sheet plus a style board, count the references you need. If it is more than ten, test the request on Sume before committing, and have a plan to merge references into a single sheet if the cap is lower than hoped.
Cost of references
References are not free either. On Sume a reference video changes the Seedance token count: the pricing code adds 15 assumed seconds and then applies a 0.6 multiplier, as described in the post on reference video cost. A request with many video references should be priced before it is sent, not after.
Sources
Related posts
More in Comparisons
- Seedance 2.5 targeted editing: vendor claim vs what Sume accepts
Together says Seedance 2.5 edits specific moments after generation. Sume lists no video_url edit for it; only gemini-omni-flash-1.1 does. The reroll workaround.
- Seedream 5.0 Lite or Nano Banana 2.1 for product shots on Sume
Seedream 5.0 Lite bills $0.04375 per image on Sume against $0.10 for Nano Banana 2.1 at 1K. Ratios, tiers, references and when the extra cost is worth paying.
- Pilot STT in 3 languages: language_code hint vs auto-detect, 30 clips
MAI-Transcribe-2-Streaming advertises 60 languages; here is a 30-clip pilot on Sume STT that compares a language_code hint with auto-detect for about 60 cents.
- Tavus PAL Maker, API, Enterprise vs the Sume avatar API: what matches
Tavus's site lists PAL Maker, a Developer API (CVI) and Enterprise. Which of them overlaps with Sume's avatar clips, and which does not. Pages read 2026-10-07.
Written by Sume