Seedance reference inputs on Sume: images, video, audio per docs
Which reference types the Seedance ids accept on /v1/videos (images, video, audio, first and last frame), how references are priced, and a working request.

The Seedance 2.x models accept image, video and audio references on Sume's /v1/videos, and the catalog lists first_frame and last_frame for image-to-video. The docs state that the Seedance 2.x models, Wan 3.0 and the MiniMax H3 models accept audio and video references, while Gemini Omni Flash 1.1 accepts video but not audio (Video generation).
Two image fields, two modes
frame_images supplies a first_frame or last_frame and starts image-to-video. input_references supplies style or content guidance and starts reference-to-video; the model treats those images as guidance, not exact frames. If you send both, frame_images wins and the request runs as image-to-video.
Request with a reference image
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "seedance-2",
"prompt": "A colossal solar flare beside a planet",
"input_references": [
{"type": "image_url", "image_url": {"url": "https://example.com/style-ref.png"}}
],
"resolution": "720p",
"duration": 8
}'Reference video changes the estimate
When a request carries a reference video, Sume's estimator adds an assumed 15 seconds of input to the billable tokens and then applies a 0.6 multiplier. For an 8 s 720p Seedance 2.5 clip that is $7.9736 with a reference video against $4.6224 without. The reservation is an estimate made at submit; read usage.cost on the finished job for the billable amount.
Check before sending
Read supported_input_references and supported_frame_images from GET /v1/videos/models for the exact id you call. A model accepts a reference type only if the field lists it.
Which models take which references
Per the docs, the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max accept audio and video references. Gemini Omni Flash 1.1, higgsfield-genjutsu and h3-max-recast accept video references but not audio. Reference images must be reachable over public HTTPS, and the video troubleshooting list names unreachable references as a cause of failed generations.
Sources
Related posts
More in Developers
- 60 video jobs in Node with a promise pool: which plans need waves
A Node promise pool for Sume video jobs, plus the fit table: 60 jobs against accepted capacity on Free, Pro, Startup and Scale, and how many waves each needs.
- Sora content rules vs Sume generation_rejected: read the job events
A prompt Sora refused may behave differently on Sume. How a rejection shows up as generation_rejected, which events to read, and why you must not assume parity.
- Sora download_content vs Sume /content: the 302 and two 409 errors
OpenAI's content endpoint took variant=video, thumbnail or spritesheet. Sume's /content redirects with a 302 and returns two different 409s. How to branch.
- Sora input_reference image: upload with uploadFile, use frame_images
OpenAI took the first-frame image as multipart input_reference. Sume wants an HTTPS URL: upload with the SDK's uploadFile, then pass it as frame_images.
Written by Sume