Seedance 2.5 reference inventory: 30 images, 10 clips, 10 tracks
A pre-flight checklist for the 30 image, 10 video and 10 audio reference slots Seedance 2.5 offers, and how to check what Sume accepts.

ByteDance Seed says Seedance 2.5 accepts up to 30 images, 10 video clips and 10 audio clips in a single pass (read 2026-10-05). Having 50 slots is not the same as needing 50 assets. Before you spend on a 30-second job, build an inventory so each asset has a reason to be there. This page is a checklist for that, plus the Sume request shape and the one place to confirm the limit your route enforces.
Sort every asset into one job
A reference is guidance, not an exact frame. Sume's video docs make the same distinction: input_references gives style or content guidance for reference-to-video, while frame_images pins first or last frames. If a picture must appear exactly as the opening frame, it belongs in frame_images, not in your reference pile.
| Slot type | Vendor ceiling | What to put there | What to leave out |
|---|---|---|---|
| Images | 30 | Product angles, the character sheet, one style card | Near-duplicates and screenshots with UI chrome |
| Video clips | 10 | Motion you want copied, a camera move, a lighting reference | Clips with watermarks or baked-in captions |
| Audio clips | 10 | A voice sample, a music bed, a sound effect | Mixed tracks with speech and music you cannot separate |
Check the clips before you attach them
Probe each video reference first. Sume's video inspect route returns probe facts, including probe.has_audio, and sampled stills, so you can see whether a clip is silent or carries speech you did not intend to pass along. Send frames: false for a probe only. Inspect and the other media routes need the file on media.sume.com; import it first with POST /v1/media-imports.
- Silent clip: do not rely on it as an audio reference.
- Clip with burned-in subtitles: the model may reproduce them; cut or crop first with video trim or video filter.
- Very long source: trim to the 3 to 8 seconds that show the move you want.
What Sume accepts
The Sume video docs say a model accepts a reference type only if its supported_input_references includes that type, and that the Seedance 2.x models accept audio and video references. They also say limits are different for each model. Fetch the catalog entry and read the capability fields rather than assuming the vendor's 30/10/10 carries over:
curl https://api.sume.com/v1/video-router/models/seedance-2.5 \
-H "Authorization: Bearer $SUME_API_KEY"A short review before you submit
Run one last pass with a second person. Ask them to look at the asset folder for two minutes and describe what the video will contain. If their description does not match your brief, the references are fighting the prompt. This is cheaper than a 30-second generation, and it catches the common failure of attaching a reference that nobody remembers adding.
Also decide in advance which assets are optional. If a first run is off, remove one reference at a time rather than rewriting the prompt, so you learn which reference was responsible.
- Count assets per slot type and compare against the catalog capability fields.
- Confirm every URL is reachable and, for media routes, already imported to
media.sume.com. - Store the list with the job id so a later run can reproduce the inputs.
Order matters less than labels
Name your files so you can match them in the prompt: product-front, product-side, voice-sample. Where a model supports positional tags, Sume documents them for Gemini Omni Flash 1.1 (<IMAGE_REF_0> and <VIDEO_REF_0>, 0-based in list order); for Seedance, describe each asset by role in plain language and keep the list order stable between takes so you can compare runs.
Sources
Related posts
More in Use cases
- Seedream 5.0 Flash style blending: a five-step brand board test plan
A five-step plan to test multi-reference style blending on Seedream 5.0 Flash, with a Sume seedream-5-lite request to run the same test on a catalog model.
- Sermon series title loop: Wan 3.0 first and last frame from one image
Make a looping sermon-series title background: send one image as both first and last frame to Wan 3.0 on Sume; 10 seconds costs $1.25 at 720p.
- Incident update as an avatar video: text first, then a 30s clip
A 30-second Sume avatar clip costs $5.52 on Standard, $7.35 on Plus or $16.50 on Max (no product). Post the text update first; the render is asynchronous.
- Shipping delay apology video with an AI avatar: script and preview
A 20-second delay notice from a Sume avatar: tone rules, a first-frame preview before the full render, and the cost on standard, plus and max.
Written by Sume