Seedance "accepts at most 12 input_references": mixes that fit

Sume returns unsupported_capability when images + videos + audio on a Seedance request exceed 12. Which mixes pass, what fal limits per type, and how to trim.

5 min readSume
All posts

A Seedance request on Sume fails with unsupported_capability and the message "seedance-2.5 accepts at most 12 input_references" when the images, videos and audio clips in input_references add up to more than 12. The fix is to drop references until the total is 12 or fewer; nine images and three videos, or nine images and three audio clips, both fit.

The check is on the total, in Sume's /v1/videos handler, and it applies to the four Seedance ids (seedance-2.5, seedance-2, seedance-2-fast, seedance-2-mini). It is a Sume limit, separate from what each provider page lists, so this post sorts out which number you are hitting.

What exactly does the 12 count?

Every entry in input_references counts once, whatever its type: image_url, video_url or audio_url. On the legacy Video Router surface the same arithmetic applies to reference_image_urls, reference_video_urls and reference_audio_urls added together. Frame images (frame_images with first_frame and last_frame) are not references; if you send frames, Sume treats the request as image-to-video and ignores input_references, per the Video generation docs.

Sume's own check does not set a separate cap for each type on Seedance. So 12 images, with no video or audio, passes Sume's check. Whether the provider accepts that is the provider's rule, and the next section shows why it differs by model.

The error comes back as a normal API error with the model name and field: "input_references", so a client can match on the code and show the message. Sume's Seedance errors post lists the other 400 and 404 responses you may see on the same endpoint.

Which mixes fit, and what does the provider say?

fal's Seedance 2.0 reference-to-video page lists up to 9 images, 3 videos and 3 audio clips, 12 files in all. Its Seedance 2.5 page lists up to 50 multimodal inputs, and ByteDance's launch post lists 30 images, 10 video clips and 10 audio clips. Sume's 12 is therefore the binding limit on 2.5, and the provider's per-type limit is the binding one on the 2.0 ids if you go past it.

Example mixes against Sume's 12-reference total, read 2026-10-02
ImagesVideosAudioTotalSume result
93012Passes the count
90312Passes the count
63312Passes the count
93315unsupported_capability: over 12
120012Passes Sume; 2.0 ids may refuse more than 9 images
130013unsupported_capability: over 12

How do you trim a request that is over 12?

Cut the references that carry the least. Duplicate angles of the same product are the first to go: three photos of one bottle rarely beat one clean photo plus a prompt tag. Audio is next, since an audio clip steers pacing, not what appears on screen. A reference video usually earns its slot, because nothing else carries motion, but it also changes the price, as covered in the price crossover post.

Order matters for the prompt. Tags such as @Image 1 and @Video 1 point at references by list position, so removing the second image renumbers the rest. Re-check the tags after trimming; the @image tag post shows how Sume normalizes casing.

What about an audio clip with no image or video?

fal's 2.0 page says audio "requires at least one image or video". Sume's /v1/videos handler, as read on 2026-10-02, does not add that rule for Seedance, but the legacy Video 1.0 docs state it for reference audio. The safe habit is to send at least one image or video alongside any audio clip, so the rule is satisfied on every surface and the question never comes up.

An audio reference also takes a slot from the 12. If you are at the limit, a voice clip is the first thing to move to a separate step, such as adding it in Timeline after the clip is generated.

Why not raise the cap to match ByteDance's 50?

That is Sume's call, and today the handler stops at 12. If you need more than 12 references, split the work: generate shots with different reference sets and join them in Timeline, or chain clips from a final frame. The 50-reference post compares the two limits in detail. Do not rely on a provider page's number as what Sume will accept; read supported_input_references from GET /v1/videos/models for the types, and the error text for the count.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume