Seedance 2.5 input_references: image, video and audio shapes
The exact JSON for an image_url, video_url and audio_url entry in input_references on POST /v1/videos, one mixed Seedance 2.5 request, and the errors you get.

Each entry in input_references on POST /v1/videos is an object with a type and a matching nested object holding a url: {"type": "image_url", "image_url": {"url": ...}}, the same with video_url, and the same with audio_url. A Seedance 2.5 request can mix all three in one array, up to 12 entries in total on Sume.
The shape follows the OpenRouter video format that Sume's Video generation docs describe, so a body written for that API carries over with a bare model id and Sume's base URL.
What does one mixed request look like?
This request gives Seedance a character photo, a motion clip and a voice clip. The prompt refers to each by position with an @Image 1 style tag, the notation ByteDance's launch post uses in its reference prompts; whether @Video 1 and @Audio 1 resolve the same way is the model's behaviour, so confirm it on a cheap run.
{
"model": "seedance-2.5",
"prompt": "@Image 1 walks along a pier, camera follows like @Video 1, she says the line in @Audio 1",
"input_references": [
{"type": "image_url", "image_url": {"url": "https://example.com/hero.png"}},
{"type": "video_url", "video_url": {"url": "https://example.com/walk.mp4"}},
{"type": "audio_url", "audio_url": {"url": "https://example.com/line.mp3"}}
],
"duration": 12,
"resolution": "720p",
"aspect_ratio": "9:16"
}Which fields are required and which are optional?
Only model and prompt are required on /v1/videos. Everything else is optional, and the types you may use in input_references depend on the model's supported_input_references.
| type | Nested object | Accepted by Seedance ids | Estimate effect |
|---|---|---|---|
| image_url | image_url.url | Yes | None |
| video_url | video_url.url | Yes | Adds 15 assumed seconds, then x0.6 |
| audio_url | audio_url.url | Yes | None |
What goes wrong, and what is the error?
A type the model does not list is refused with unsupported_capability, and the message names the field. A fourth type, say file_url, is not in the schema. Seedance totals above 12 are refused with "accepts at most 12 input_references", covered in the 12-reference post. Duplicate fields are the other trap: if the body also carries frame_images, the references are dropped and the job runs as image-to-video.
Another trap is the field name. input_references is the /v1/videos name. The Video Router surface at /v1/video-router/generate uses reference_image_urls, reference_video_urls and reference_audio_urls as flat arrays of strings. Mixing the two shapes in one body will not do what you want.
Test the whole body at 480p first. A 480p, 4-second request is the shortest, lowest-resolution setting in the catalog's seedance-2.5 range, and it will show whether your URLs resolve and your tags point at the right entries before you pay for 720p or 1080p.
What is the same request on the Video Router path?
On the legacy /v1/video-router/generate path the same three references are flat arrays of URL strings. The model, the 12-entry total and the pricing are the same; only the field names change.
{
"model": "seedance-2.5",
"prompt": "@Image 1 walks along a pier, camera follows like @Video 1",
"reference_image_urls": ["https://example.com/hero.png"],
"reference_video_urls": ["https://example.com/walk.mp4"],
"reference_audio_urls": ["https://example.com/line.mp3"],
"duration": 12,
"resolution": "720p",
"aspect_ratio": "9:16",
"mode": "async"
}What should the media URLs look like?
Use public HTTPS URLs that the provider can fetch without a login. Sume's docs say a failed generation can come from reference images that are not accessible over public HTTPS or are in an unsupported format, and the API reference says generation requests accept fetchable public HTTPS media URLs directly.
ByteDance's launch post lists up to 30 images, 10 video clips and 10 audio clips for Seedance 2.5, and the Dreamina launch release says the audio files can be music and dialogue. How closely a clip follows an audio reference is a model matter, not a Sume setting, and the audio reference post covers what to expect.
How do you choose between the three types?
Images say who and what: a character, a product, a scene. Video says how it moves: a camera path or an action, which is why a video reference is the only type that raises the estimate, as in the camera-move post. Audio says how it sounds and paces: a voice line or a beat. Start with images alone, add a video only if motion is wrong, and add audio last.
When a result is wrong, change one thing at a time. If the character drifts, add or sharpen an image. If the camera does nothing, swap in a clearer video reference. If the voice is off, shorten the audio clip so it covers only the line. Each change is one entry in the array, and each keeps you inside the 12-entry total.
Sources
Related posts
More in Developers
- Seedance 2.5 reference images in Python with asyncio and httpx
A runnable Python script: send reference images to seedance-2.5 on Sume's /v1/videos, poll every 30 s with asyncio, save the MP4, and handle the errors.
- Seedance 2.5 reference images in TypeScript with Node fetch
A Node 18+ ESM script: submit reference images to seedance-2.5 on Sume's /v1/videos, poll every 30 s with fetch, print the video URL. Rules and mistakes.
- Seedance "accepts at most 12 input_references": mixes that fit
Sume returns unsupported_capability when images + videos + audio on a Seedance request exceed 12. Which mixes pass, what fal limits per type, and how to trim.
- One image to Seedance: reference, or first frame?
On Sume a single reference image with no frame field is priced and routed as reference-to-video; add a first frame to get image-to-video. How to choose.
Written by Sume