Seedance 2.0 reference limits: 9 images, 3 videos, 3 audio?
fal lists Seedance 2.0 reference-to-video at up to 9 images, 3 videos and 3 audio clips, 12 files in all. On Sume the API enforces the 12-file total. Details.

fal's Seedance 2.0 reference-to-video page lists up to 9 images, 3 videos and 3 audio clips, with no more than 12 files across all types. On Sume, a request to seedance-2, seedance-2-fast or seedance-2-mini is refused with 400 unsupported_capability when it carries more than 12 input_references in total; the API code checks that total for these ids and no per-type count.
The vendor numbers below are from fal's page, read 2026-09-29. The Sume behavior comes from the Video generation docs and the API's request checks.
What limits does fal state for Seedance 2.0 references?
Treat these as the model host's limits for files. They are the numbers to plan your assets against.
| Type | Count | Formats | Stated constraints |
|---|---|---|---|
| Images | Up to 9 | JPEG, PNG, WebP | Max 30 MB each |
| Videos | Up to 3 | MP4, MOV | Combined duration 2–15 s, total under 50 MB |
| Audio | Up to 3 | MP3, WAV | Combined duration 15 s or less, max 15 MB each; needs at least one image or video |
| All types | 12 | Total files must not exceed 12 |
What does Sume check on a Seedance 2.0 request?
Sume accepts the types the model row lists: for the Seedance 2.x ids that is image_url, video_url and audio_url. It then counts every reference and refuses the request past 12. The docs publish no file-size, format or clip-length limits for references, so use fal's numbers above as the planning limits.
Sume's count is a total only. A request with 10 images and no other references passes that count, but it is over the 9-image limit fal states, so stay at 9 images, 3 videos and 3 audio clips.
How do I send images, a video and audio together?
Each entry has a type and a matching object with a public HTTPS url. Put the images, video and audio in one input_references array.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "seedance-2",
"prompt": "The person from the first image walks through the street in the video",
"duration": 8,
"resolution": "720p",
"aspect_ratio": "16:9",
"input_references": [
{"type": "image_url", "image_url": {"url": "https://example.com/person.png"}},
{"type": "video_url", "video_url": {"url": "https://example.com/street.mp4"}},
{"type": "audio_url", "audio_url": {"url": "https://example.com/ambience.mp3"}}
]
}'What happens with 13 references?
The create call returns 400 unsupported_capability with the message "seedance-2 accepts at most 12 input_references." (the model name is the one you sent). Trim the array and send a new request; resending the same body gets the same refusal. Choose the references that matter most: keep the image that fixes the subject, one video for motion and one audio clip for sound, and drop near-duplicate images first.
Audio needs care, because a soundtrack alone is not enough. fal states that audio references require at least one image or video, so send audio alongside one of them.
Is the limit the same for Seedance 2.5?
The 12-file total on Sume is the same for seedance-2.5. This post's per-type numbers are fal's Seedance 2.0 page; Seedance 2.5 references covers 2.5 separately.
Sources
Related posts
More in Developers
- Seedance 2.5 API key: get one and send a first request
Create a Sume API key, POST to /v1/videos with model seedance-2.5, poll the job and download the file. One curl request, the auth rules and the fields to set.
- How long does a Seedance 2.5 video take, and how do you wait for it?
Sume's docs say video generation takes 30 seconds to several minutes. Poll GET /v1/videos/{id} every 30 seconds or pass callback_url for a signed webhook.
- BytePlus ModelArk Seedance 2.5: limits, tasks and how Sume maps
What BytePlus's ModelArk tutorial documents for Seedance 2.5 (task types, drafts, rate limits, retention) and which concepts exist in Sume's video API.
- C# speech to text: transcribe audio files with HttpClient
Speech to text in C#: POST the audio file's URL with HttpClient, poll the job, then read the transcript and word timestamps from the JSON result.
Written by Sume