Give every Seedance 2.5 reference a role in the prompt
With up to 12 references per Sume request, say in the prompt what each one is for. Roles for images, videos and audio, and a worked reference-to-video example.

In a Seedance 2.5 reference-to-video request, tell the model what each reference is for: an image for identity, another for place, a video for motion or camera, an audio file for voice or music. Sume does not define a tag syntax for Seedance in the docs read on 2026-10-04, so write the roles in plain words in the prompt and keep the reference list short.
Why roles matter
A bag of references is ambiguous. Three photos might be three characters, three views of one character, or three locations. ByteDance describes references as a way to control performance, lighting, shadow and camera movement (ByteDance Seed: Seedance 2.0), which are different jobs for different inputs. Naming the job reduces the ways the model can misread your input.
A role table
The fields below are the Sume Video Router fields. The "role" column is a prompt-writing convention, not an API setting.
| Field | Typical role | Sume cap |
|---|---|---|
| reference_image_urls | identity of a person or product; a place; a style | 9 |
| reference_video_urls | camera move or motion to follow | 3 |
| reference_audio_urls | music or voice to pace the clip | 3, needs an image or video |
| all together | any mix of the three | 12 |
Say it in the prompt
Refer to inputs by what they are, and keep the order of the arrays predictable. The Sume docs define positional tags such as <IMAGE_REF_0> only for Gemini Omni Flash 1.1 in the Video Router docs; do not assume the same syntax works for Seedance. Describe it in words instead.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: roles-001" \
-d '{
"model": "seedance-2.5",
"prompt": "The woman in the first image is the lead. The second image is the cafe where the scene takes place. Follow the slow dolly-in of the reference video. Pace the cuts to the reference music.",
"reference_image_urls": ["https://example.com/lead.png","https://example.com/cafe.png"],
"reference_video_urls": ["https://example.com/dolly.mp4"],
"reference_audio_urls": ["https://example.com/theme.mp3"],
"resolution": "480p",
"duration": 10,
"mode": "async"
}'Mind the price of video references
A reference video changes the bill. In Sume's pricing code a request with a reference video is priced on the output length plus an assumed 15 seconds of input, at 0.6 of the base, so short outputs cost more and 30-second outputs cost less than the no-video price. At 720p, a 10-second output goes from $5.78 to about $8.67, and a 30-second output goes from $17.33 to about $15.60. Confirm in the live catalog before relying on this.
Test with one image role at a time first. If the identity is wrong with images alone, adding a video will not fix it.
Sources
Related posts
More in Developers
- A Go client for Sume from OpenAPI, with Retry-After
The docs list only a TypeScript SDK. Generate a Go client from the live OpenAPI schema, send x-api-key, and back off on 429 with retry-after.
- Asset Studio's four-step Omni workflow, rebuilt as Sume API calls
Asset Studio's flow is anchor, concepts, refine, export. Here is each step as a Sume call: /v1/videos, video-trim, timeline fit modes and a cropped square.
- Google's June 15 deprecation notice gave 15 and 63 days: run a drill
Google announced Veo and Imagen 4 deprecations on Jun 15, 2026 with shutdowns Jun 30 and Aug 17. Here is a five-step drill that fits inside the shorter window.
- GPT-6.1 Sol function tool for Sume runs: clamp the cap in code
A Responses API function tool that lets GPT-6.1 Sol start a Sume Agent Completion: call_id as Idempotency-Key and a spend cap the model cannot exceed.
Written by Sume