Seedance 2.5 input_references: image, video and audio shapes

The exact JSON for an image_url, video_url and audio_url entry in input_references on POST /v1/videos, one mixed Seedance 2.5 request, and the errors you get.

5 min readSume
All posts

Each entry in input_references on POST /v1/videos is an object with a type and a matching nested object holding a url: {"type": "image_url", "image_url": {"url": ...}}, the same with video_url, and the same with audio_url. A Seedance 2.5 request can mix all three in one array, up to 12 entries in total on Sume.

The shape follows the OpenRouter video format that Sume's Video generation docs describe, so a body written for that API carries over with a bare model id and Sume's base URL.

What does one mixed request look like?

This request gives Seedance a character photo, a motion clip and a voice clip. The prompt refers to each by position with an @Image 1 style tag, the notation ByteDance's launch post uses in its reference prompts; whether @Video 1 and @Audio 1 resolve the same way is the model's behaviour, so confirm it on a cheap run.

{
  "model": "seedance-2.5",
  "prompt": "@Image 1 walks along a pier, camera follows like @Video 1, she says the line in @Audio 1",
  "input_references": [
    {"type": "image_url", "image_url": {"url": "https://example.com/hero.png"}},
    {"type": "video_url", "video_url": {"url": "https://example.com/walk.mp4"}},
    {"type": "audio_url", "audio_url": {"url": "https://example.com/line.mp3"}}
  ],
  "duration": 12,
  "resolution": "720p",
  "aspect_ratio": "9:16"
}

Which fields are required and which are optional?

Only model and prompt are required on /v1/videos. Everything else is optional, and the types you may use in input_references depend on the model's supported_input_references.

input_references entry shapes on /v1/videos, read 2026-10-02
typeNested objectAccepted by Seedance idsEstimate effect
image_urlimage_url.urlYesNone
video_urlvideo_url.urlYesAdds 15 assumed seconds, then x0.6
audio_urlaudio_url.urlYesNone

What goes wrong, and what is the error?

A type the model does not list is refused with unsupported_capability, and the message names the field. A fourth type, say file_url, is not in the schema. Seedance totals above 12 are refused with "accepts at most 12 input_references", covered in the 12-reference post. Duplicate fields are the other trap: if the body also carries frame_images, the references are dropped and the job runs as image-to-video.

Another trap is the field name. input_references is the /v1/videos name. The Video Router surface at /v1/video-router/generate uses reference_image_urls, reference_video_urls and reference_audio_urls as flat arrays of strings. Mixing the two shapes in one body will not do what you want.

Test the whole body at 480p first. A 480p, 4-second request is the shortest, lowest-resolution setting in the catalog's seedance-2.5 range, and it will show whether your URLs resolve and your tags point at the right entries before you pay for 720p or 1080p.

What is the same request on the Video Router path?

On the legacy /v1/video-router/generate path the same three references are flat arrays of URL strings. The model, the 12-entry total and the pricing are the same; only the field names change.

{
  "model": "seedance-2.5",
  "prompt": "@Image 1 walks along a pier, camera follows like @Video 1",
  "reference_image_urls": ["https://example.com/hero.png"],
  "reference_video_urls": ["https://example.com/walk.mp4"],
  "reference_audio_urls": ["https://example.com/line.mp3"],
  "duration": 12,
  "resolution": "720p",
  "aspect_ratio": "9:16",
  "mode": "async"
}

What should the media URLs look like?

Use public HTTPS URLs that the provider can fetch without a login. Sume's docs say a failed generation can come from reference images that are not accessible over public HTTPS or are in an unsupported format, and the API reference says generation requests accept fetchable public HTTPS media URLs directly.

ByteDance's launch post lists up to 30 images, 10 video clips and 10 audio clips for Seedance 2.5, and the Dreamina launch release says the audio files can be music and dialogue. How closely a clip follows an audio reference is a model matter, not a Sume setting, and the audio reference post covers what to expect.

How do you choose between the three types?

Images say who and what: a character, a product, a scene. Video says how it moves: a camera path or an action, which is why a video reference is the only type that raises the estimate, as in the camera-move post. Audio says how it sounds and paces: a voice line or a beat. Start with images alone, add a video only if motion is wrong, and add audio last.

When a result is wrong, change one thing at a time. If the character drifts, add or sharpen an image. If the camera does nothing, swap in a clearer video reference. If the voice is off, shorten the audio clip so it covers only the line. Each change is one entry in the array, and each keeps you inside the 12-entry total.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume