ElevenLabs sound effects: prompt influence, 48 kHz WAV, and Sume
ElevenLabs SFX offers high or low prompt influence and 48 kHz WAV. Sume has no sound-effects route; a Music prompt plus a split makes a sting. What you lose.

ElevenLabs sound effects let you choose how literally the model follows your prompt (high or low prompt influence) and return 48 kHz WAV for non-looping effects. Sume has no dedicated sound-effects route, so the closest documented substitute is a short Music generation, optionally cut with a split job, and it gives you neither a prompt-influence setting nor a guaranteed effect-style result.
The ElevenLabs details come from its sound effects documentation (read 2026-10-10). The Sume side comes from the Music 1.0 and Timeline audio pages. I could not find a sound-effects endpoint in Sume's docs or OpenAPI paths, which list music, TTS, STT, audio detach and timeline audio routes for audio.
What the ElevenLabs page lists
The page gives a duration range of 0.1 to 30 seconds, with automatic duration as the default; a loop option for effects longer than 30 seconds; two prompt-influence settings, where high is a more literal interpretation and low is more creative with added variations; MP3 for all effects; and WAV at 48 kHz for non-looping effects. It states a cost of 40 credits per second when a duration is specified.
| Control | ElevenLabs sound effects | Sume closest route |
|---|---|---|
| Dedicated SFX route | Yes | None documented |
| Length | 0.1 to 30 seconds | Music has no duration field; cut with timeline audio split |
| Literal vs creative | High or low prompt influence | Write the literalness into the prompt |
| Output | MP3; 48 kHz WAV for non-looping | Audio artifact, usually audio/mpeg; split output defaults to WAV |
| Price shape | 40 credits per second when duration is set | $0.125 per Music generation plus $0.01 per split job |
A Music prompt as a sting
Sume's Music brief works for stings, risers and hits that are musical in nature. Write a precise brief with a tempo number, one or two instruments, one named moment, and end with 'Instrumental, no vocals.' Ask for a short track in the prompt, for example 'a 6-second riser'. The Music page says a non-empty negative_prompt is rejected, so put exclusions such as 'no spoken word' inside the positive prompt.
This works poorly for non-musical sounds. A door slam or a footstep is not a musical brief, and Sume's docs make no claim that the Music engine produces literal foley. Treat any such result as an experiment you check by ear, at $0.125 per try.
Cutting the effect out
If the generation contains the hit you want plus extra tail, timeline audio can slice it. Send operation: "split", the Sume-hosted file URL, and up to 20 ranges; each range returns its own audio_url and the job costs $0.01. WAV is the default output and is sample-exact, while MP3 adds priming padding at every edge, so keep WAV for short hits you will place against picture.
The Music artifact already sits on media.sume.com, so it should satisfy the Sume-hosted requirement; the docs say off-host URLs are rejected at admission and you import first.
When to use something else
If you need literal foley at 48 kHz with a prompt-influence slider, ElevenLabs lists exactly that, and Sume does not. If the sound is already in a video you generated, there may be no need to synthesize it: audio detach can pull the track out of a Sume-hosted clip for $0.01, and a range of it can become the sting. See the Sume SFX generator page for the broader workflow.
A fair summary: Sume covers the music and mixing half of a sound design, and does not claim the effects half. Say that plainly in your own project docs so no one budgets a foley step that does not exist.
Sources
Related posts
More in Comparisons
- Gemini TTS has 30 prebuilt voices; Sume TTS takes avatar voice ids
Google's Gemini TTS page lists 30 prebuilt voices. Sume TTS takes a voice UUID or voi_ id, or an avatar whose voice.status is ready; cloning stays app-only.
- Grok Imagine's 4 keyframes and 7 references vs Sume's one image
xAI's Grok Imagine 1.5 takes up to 4 keyframes and up to 7 references. Sume's grok-imagine-video-1.5 row takes one image only. Rows to use for multi-image work.
- Grok Imagine Lite upscales to 1080p: what Sume offers instead
xAI describes Grok Imagine Video 1.5 Lite as lowest cost with upscaled 1080p. Sume has no Lite row; it lists grok-imagine-video-1.5 at a flat rate. Compare.
- Grok Imagine's 3 voice references vs Sume reference-audio rows
xAI's Grok Imagine 1.5 takes up to 3 voice references. Sume's Grok row takes none; Seedance, Wan 3.0 and MiniMax accept reference audio under limits.
Written by Sume