Sony Woosh sound effects model: does Sume have SFX or video-to-audio?
Sony AI's Woosh makes sound effects and video-to-audio. Sume has no SFX route; it offers video models with native audio, plus music stings. What to use instead.

Sume does not carry Sony's Woosh and has no sound-effects or video-to-audio route in its API. If you need effects for a clip, the options on Sume are video models that generate their own synced audio, a short music generation for a sting, or sound you bring yourself.
Sony AI announced Woosh on May 18, 2026; the post was read on 2026-10-10. The facts below are from that post and from the Sume repository, and nothing here claims Sume can reproduce what Woosh does.
What Sony says Woosh is
Sony AI describes Woosh as a sound-effect foundation model. The post separates two releases: a private model trained on licensed libraries (Pro Sound Effects and BOOM are named), and a public open-weights model for research and non-commercial use. It supports video-to-audio, where a clip is the input and a text prompt is optional.
The post does not state output sample rate or clip length, and the public weights are non-commercial. For a shipped ad or product video, that second point matters: the openly downloadable version is not the one to build a paid workflow on.
What Sume actually lists for sound
Sume's audio routes in the OpenAPI file are TTS, STT, music, audio detach, timeline audio and lip-sync. There is no effects-generation endpoint, so a prompt like footsteps on gravel cannot be sent to a sound model. Sume does have two paths that put sound on a clip.
First, some video models return audio with the picture. The API schema marks gemini-omni-flash-1.1 as always producing native synced audio; the generate_audio flag is accepted there but changes nothing. The MiniMax H3 and H3 Max models always produce native stereo audio, and generate_audio set to false is rejected for them. See Videos for each model's limits; Omni Flash 1.1 accepts 3 to 10 seconds per clip.
Second, if you generate a clip and only want its sound, audio detach pulls the track as wav or mp3 for $0.01 per job, and a range lets you take just the part you want, as in a three-second sting from an Omni clip.
| Route | Prompt for effects? | Commercial use | Where |
|---|---|---|---|
| Woosh private model | Text and optional video | Per Sony's terms; not stated in the post | Not on Sume |
| Woosh public weights | Text and optional video | Research and non-commercial | Not on Sume |
| Omni Flash 1.1 native audio | Describe the sound in the video prompt | Check Sume terms | Sume video |
| MiniMax H3 / H3 Max native audio | Describe the sound in the video prompt | Check Sume terms | Sume video |
| Music generation as a sting | Describe a short hit | Check Sume terms | Sume music, $0.125 |
Using a music sting as a stand-in
Music 1.0 is not an effects model, but a short prompt can produce a riser, hit or whoosh-style transition. It costs a flat $0.125 per generation and there is no duration parameter, so you ask for a short piece in the prompt and trim it afterward. Expect music, not a Foley recording.
Where a real, recorded effect matters, such as a door close that must match a specific frame, a licensed library you control is the safer way, and Sume's Timeline can then place your file. See AI sound effects for video without a SFX API for the same workaround in more detail.
Which route fits which job
The choice is mostly about whether the sound must be tied to a picture you are already generating.
- A generated clip needing ambient and action sound together: use a video model with native audio and describe the sound.
- A custom hit or transition: try a one-line music prompt, then trim.
- Frame-exact effects on existing footage: bring licensed effects; Sume has no video-to-audio route.
- Research on sound-effect models: Sony's open weights fit, if non-commercial use is acceptable.
What to check before relying on this
Native audio is a side effect of video generation, not a controllable effects pipeline, so listen to every output. Re-read Sony's page for license changes, and check Sume's catalog for new audio routes before you plan around a gap that may have closed.
Sources
Related posts
More in Comparisons
- Spreadsheet rows to videos: Creatomate or Sume bulk runs?
Creatomate ties a template to a spreadsheet; Sume bulk runs queue up to 100 Format runs at concurrency 1 to 16. How rows map, and how many queues 250 rows take.
- Suno Italy probe: four terms clauses to check before ad music
Italy's AGCM opened a probe into Suno's terms on Oct 6, 2026. The four flagged clauses, and what to check in any AI music tool before putting a track in an ad.
- Suno Speech makes voice and music in one pass. What does Sume do?
Suno's Speech beta (Oct 1, 2026) makes voice and music as one track. Sume has no such model; it layers TTS, Music and a Timeline soundtrack. Trade-offs inside.
- Suno v6-mini for all users vs Sume Music at $0.125 a track
Suno v6-mini is open to every user, while v6 and v6-wild need Pro or Premier. Sume does not list Suno; it sells one Lyria track for a flat $0.125.
Written by Sume