Image-to-video with sound by API: Sume ids that take a first frame

Kandinsky 6.0 has an image-to-audio-video mode. On Sume these ids take a first_frame image and return sound, and these do not. Request code included.

5 min readSume
All posts

Send frame_images with a first_frame entry to POST /v1/videos and pick an id whose catalog row reports audio, such as seedance-2-fast, wan-3.0 or gemini-omni-flash-1.1. That is Sume's equivalent of the image-to-audio-video mode in Kandinsky 6.0, which the kandinsky-6 repository lists next to text-to-audio-video (read 2026-10-11).

Not every image-to-video id makes sound, though, and one popular id needs an image to work at all. The table below separates them.

Which ids take a first frame and make sound

Sume's video docs say frame_images entries carry a frame_type of first_frame or last_frame, and that a model accepts only the frame types in its supported_frame_images. This table is from the catalog code on the main branch, read on 2026-10-11.

Image-to-video frame support and audio by id (catalog code, read 2026-10-11)
Sume idFrames acceptedMakes audioDurations
seedance-2-mini, -fast, seedance-2first and lastyes4-15 s
seedance-2.5first and lastyes4-30 s
kling-3first and lastyes4-15 s
wan-3.0first and lastyes2-30 s
minimax-h3, minimax-h3-maxfirst and lastyes5-15 s
gemini-omni-flash-1.1first and lastalways on3-10 s
grok-imagine-video-1.5first onlyno4-15 s

A request that returns a clip with sound

This script follows the polling flow in the docs: submit, wait 30 seconds between polls, then read unsigned_urls. Replace the image URL with a public HTTPS image. Per the docs, Sume reserves the provider list price times 1.25 at submit, so your balance must cover that before the job starts.

import os
import time
import requests

H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
body = {
    "model": "seedance-2-fast",
    "prompt": "The kettle steams; a quiet kitchen with soft room sound",
    "duration": 5,
    "resolution": "720p",
    "generate_audio": True,
    "frame_images": [{
        "type": "image_url",
        "image_url": {"url": "https://example.com/first-frame.png"},
        "frame_type": "first_frame",
    }],
}
r = requests.post("https://api.sume.com/v1/videos", headers=H, json=body, timeout=60)
r.raise_for_status()
poll = r.json()["polling_url"]
while True:
    time.sleep(30)
    s = requests.get(poll, headers=H, timeout=60).json()
    if s["status"] == "completed":
        print(s["unsigned_urls"][0])
        break
    if s["status"] in ("failed", "cancelled"):
        print(s.get("error"))
        break

Three details that cause rejected requests

  • If you send both frame_images and input_references, the docs say frame_images controls the mode and Sume treats the request as image-to-video.
  • size returns a 400 because each v1 model reports supported_sizes: null; use resolution and aspect_ratio instead.
  • For grok-imagine-video-1.5 the request schema requires one input image and rejects aspect_ratio, an end frame and generate_audio. Use a different id when you want sound or an end frame.
  • For gemini-omni-flash-1.1 the audio cannot be turned off, so a silent clip needs another id.

Picking one

For a quick, cheap test of how a still moves with sound, start at seedance-2-mini or seedance-2-fast at 720p or lower. For a clip longer than 15 seconds from one image, seedance-2.5 and wan-3.0 accept up to 30. Pin the id while you compare, because sume/auto hides which family served the request and would blur the comparison.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume