AI sound effects: Suno Sounds vs a music prompt
Suno Sounds V5.5 generates sound effects for 2 credits flat. Sume has no SFX endpoint; a Music Router prompt with no vocals gives ambient audio you must check.

Suno Sounds V5.5 is a sound generation model priced at 2 credits flat, according to a Versely roundup. Sume has no dedicated sound-effects model. The closest route is the Music Router with a prompt that asks for no vocals; it returns audio, but the docs promise a music engine, not effects, so listen before you ship.
What Versely reports
From the Versely state of AI voice post, read on 2026-10-03:
| Item | Detail |
|---|---|
| Release | March 26, 2026 |
| Price | 2 credits, flat |
| Type | Sound generation, not speech synthesis |
What Sume offers
Sume documents music generation through the Music Router, where sume/music-auto resolves to Lyria 3.5 today. Speech is a separate tool (tts_create on MCP). There is no sound-effects model in the docs, so any effect you get comes from steering a music engine with text.
A prompt that asks for ambience
The Music prompt guidance says to close with one clause such as "Instrumental, no vocals." and to put exclusions in the positive prompt, since a non-empty negative_prompt is unsupported. Length is steered in the prompt because duration and duration_seconds are rejected. Provider lyrics metadata is model-reported, not an audio measurement.
curl -X POST https://api.sume.com/v1/music-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: sfx-door-001" \
-d '{
"model": "sume/music-auto",
"prompt": "A 10-second ambient texture: a heavy wooden door creaking open in an empty hall, low room tone, no melody. Instrumental, no vocals."
}'How the two differ in practice
A sound model is trained to produce events such as hits, whooshes or footsteps. A music engine is trained to produce songs, so the output may add rhythm or harmony you did not ask for. The docs call the music brief axes creative directions, not guaranteed output settings, and tell you to verify the generated audio.
- Generate two or three takes and keep the one that is closest to a plain effect.
- Trim with the video and audio tools you already use, or lay the file under footage in Timeline.
- If you need a specific library sound, use a source built for effects and bring the file in as an audio input.
Sources
Related posts
More in Models
- AI video releases and shutdowns in 2026: one dated timeline
Twelve dated events from Google, Luma, Runway, OpenAI and ElevenLabs, January to October 2026, each taken from a vendor page read on 2026-10-03.
- Eleven v4 stacked tags: direct emotion without them
ElevenLabs v4 adds stackable expression tags and 10-second cloning. Sume's docs list no clone route; here is how to direct a voice via tts_create.
- eleven_v4 vs eleven_v4_turbo: model IDs, endpoints, which to pick
ElevenLabs lists eleven_v4 for expressive speech with cloning in 90+ languages and eleven_v4_turbo at about 100 ms median latency. Which fits a video pipeline.
- ElevenLabs languages: Flash v2.5 has 32, Multilingual v2 29, v4 90+
ElevenLabs lists 32 languages for Flash v2.5, 29 for Multilingual v2 and 90+ for v4 and v4 Turbo. Check your markets against the model, then log it per job.
Written by Sume