Eleven v4 tags like [light rain] vs a separate sound bed on Sume
Eleven v4 puts effects such as [light rain] inside the speech. Sume keeps voice and bed separate, with gain_db, loop and duck_db. The trade-off for editing.

ElevenLabs' Eleven v4 announcement says you can direct delivery with inline tags such as [laughs], [said angrily in French accent], [light rain] and [phone buzzing] (read 2026-10-05). The last two are not speech: they ask the model to add ambience or an effect inside the same audio file.
The trade-off
With one file you get a finished take in one call. The cost is that the rain, the buzz and the voice are baked together. If the client wants the rain quieter, or the buzz two seconds later, you regenerate the line. With separate stems you change one number.
The Sume way: voice and bed apart
Sume TTS returns a voiceover only, with no inline effect tags. Ambience or music is a second file that you import or generate, then place in a Timeline render as the soundtrack. The documented soundtrack fields are url, gain_db, loop, fade_out_seconds (up to 10) and duck_db (0 to 20), and ducking needs a real audio spine under it, not silence.
gain_dbsets the bed level.looprepeats a short rain loop under a long voiceover.fade_out_secondsends the bed cleanly.duck_dblowers the bed while the voice speaks.
Which to pick
| Need | Inline tag in speech | Separate bed on a timeline |
|---|---|---|
| One quick take | Best: one call | Two files and a render |
| Change effect volume later | Regenerate the line | Edit gain_db |
| Reuse the same rain on 20 videos | Re-prompt each time | Import once, reference everywhere |
| Effect timed to a word | The model decides | You place it |
A budget example
A 600-character voiceover costs $0.0285 on Sume TTS, billed as 3 cents. Timeline renders are priced separately; check the Timeline docs for the current rate before you plan a batch. The point is that the bed is reusable: pay once for a rain loop and put it under every video.
Sources
Related posts
More in Comparisons
- Directing delivery: Eleven v4 audio tags vs Gemini TTS style
Eleven v4 puts direction inline as audio tags like [laughs]; Gemini TTS adds a separate style field. Sume sends transcript, voice and language only.
- Eleven v4 IPA support vs Sume's pronunciation_dict_id for brand names
ElevenLabs says Eleven v4 improves IPA phoneme support. Sume TTS has an optional pronunciation_dict_id. How each helps with brand names today.
- Voice agent TTS latency: Eleven v4 Turbo 150ms vs Cartesia Sonic
Eleven v4 Turbo claims about 150ms to first speech; Cartesia Sonic-3.6 claims under 90ms. Not like-for-like; Sume TTS is not for live agents.
- Eleven v4 or Sonic 3.6 for ad voiceover through an API
Eleven v4 lists $0.022 per 1K characters, Sonic 3.6 $0.038 at list. Sume serves Sonic 3.6 only, at list x1.25. Here is how to pick for ad reads.
Written by Sume