Pika Soundtrack: video in, sound out. What Sume has instead
Pika Soundtrack scores a finished video. Sume's docs offer generate_audio at video creation and a Timeline soundtrack bed, but no video-to-sound endpoint.

Pika Soundtrack takes an existing video and generates sound effects, voice, music and ambience for it. The Sume docs read 2026-10-01 describe no endpoint that does this. What they do have is a generate_audio flag on video creation and a Timeline soundtrack bed you lay under a finished cut.
What does Pika Soundtrack do?
Per Pika's August 18 post, you upload a video and it generates motion-aware sound effects, voice, music and ambience that follow the action, with an optional instruction about what to emphasize, include or leave out. It says the model is live in the Pika API Club. Pika's speed and benchmark statements are its own and are not repeated here.
What does Sume offer for sound on video?
Two separate things, neither of which reads a finished video and invents sound for it. See Video generation and Timeline 1.0.
| Option | When it applies | What the docs say |
|---|---|---|
generate_audio | At video generation | Whether to generate audio alongside the video; defaults to the model's capability |
| Gemini Omni Flash 1.1 | At video generation | Native audio is always on |
Timeline soundtrack | Joining finished clips | Optional bed: url, gain_db, loop, fade_out_seconds up to 10, duck_db |
| Ducking | Soundtrack under a voice spine | duck_db needs a real audio spine; with silence the request fails duck_requires_audio_spine |
How do I score a silent clip on Sume?
Make the sound as its own file, then place it. Generate music or voice with the audio endpoints, import it as Sume-hosted media, and give it to Timeline as the audio.url spine or the soundtrack bed. Timeline renders at $0.10 per output minute. For sound effects specifically, read sound effects for video without a SFX API.
Will the sound follow the action?
Not automatically. Timeline places the bed you supply; it does not analyze the picture to time impacts. Matching sound to motion is Pika Soundtrack's stated purpose, and the Sume docs make no equivalent claim. If sync matters, generate with a model that has audio on, or time your own audio against the cut.
Sources
Related posts
More in Models
- 30-second video with sound: Pika Video Studio vs Sume ids
Pika Video Studio makes up to 30 seconds with sound. On Sume, seedance-2.5 and wan-3.0 reach 30 seconds; every other catalog model stops at 15.
- Pika Video Studio's 50 references vs Sume's per-model limits
Pika Video Studio attaches up to 50 references. Sume's caps are per model: 10 images and 3 short clips on Omni Flash, 1 to 8 images on Genjutsu.
- Qwen-Audio-3.0-TTS languages: 16, and Sume's language field
Qwen-Audio-3.0-TTS lists 16 languages. On Sume TTS you set the language field to a BCP-47 or ISO-639 code for every non-English transcript.
- Qwen-Image 2.1 ControlNet Union Fun in ComfyUI vs reference images
ComfyUI added Qwen-Image 2.1 ControlNet Union Fun for editing on 2026-09-29. Sume has no ControlNet input; it edits with reference images and masks.
Written by Sume