ElevenLabs Music stems: Creator 2 and 4, Pro 6, and the Sume route
ElevenLabs lists 2 and 4 stems on Creator and up to 6 on Pro. When a video editor needs stems, and when a Sume duck_db setting does the job.

ElevenLabs lists stems as a paid-tier feature of its music app: 2 and 4 stems on Creator, up to 6 on Pro. Sume's Music Router docs describe one mixed audio file and no stem output, so if you need to remix drums or vocals separately you need a vendor that exports stems. If you only need the music to sit under a voiceover, a Timeline render with duck_db does that without stems.
What ElevenLabs says
The ElevenLabs music page lists stems on Creator (2 and 4) and Pro (up to 6). The music page also mentions "stem editing" under its API section, but the API docs page we read does not mention stems, so confirm in the API reference that stems are exposed before you build on them. That distinction matters to a pipeline builder: a feature on a pricing grid is not a feature on an endpoint.
| Question | ElevenLabs | Sume Music Router |
|---|---|---|
| Stems | Creator: 2 and 4; Pro: up to 6 (music page) | None; one audio artifact |
| Stems in API docs page | Not mentioned on the capabilities page | Not applicable |
| Output | MP3 or WAV per API docs | One audio artifact in result.artifacts[] |
| Mix under a voice | Done in your editor | Timeline soundtrack with duck_db 0 to 20 |
Three jobs people call "stems"
Most requests fall into one of three jobs, and only one of them truly needs separated tracks.
- Quieter music under speech: use the Timeline
soundtrack.duck_db(0 to 20) andgain_db. No stems required. - A version without drums for a calm cut: write it in the brief for a second take on Music Router, or take stems from a vendor that exports them.
- A karaoke or vocal-only file: that needs separation or a generator that outputs it. Sume has no such step.
The Sume path in practice
Generate with POST /v1/music-router/generate, then place the artifact as soundtrack.url in a Timeline 1.0 render. The soundtrack fields are url, gain_db, loop, fade_out_seconds (up to 10) and duck_db; duck_db needs a real audio spine, so it works when a voiceover is the spine and fails with duck_requires_audio_spine when the spine is silence.
To change a track's length first, use timeline audio: split slices up to 20 ranges, each a new durable file, for $0.01 per job.
Be honest about the gap
If stems are a hard requirement for your workflow, Sume's Music Router docs describe no way to meet it and this post will not talk around that. If the requirement is really "voice must stay clear", ducking is the shorter route. Verify output from either path by listening before it ships.
Pricing and tier names move; check the live pages before you buy a plan for stems alone.
Sources
Related posts
More in Models
- ElevenLabs Music v2.5 genre strengths and how to test them
ElevenLabs says v2.5 is strongest on vocal-led and acoustic-heavy genres such as R&B, soul, rock and orchestral. Test your genre with three briefs.
- Gemini 2.5 Flash Image shut down Oct 2: what to call on Sume instead
Google's gemini-2.5-flash-image, the original Nano Banana, was set to shut down Oct 2, 2026. Which Sume image ids to test as a replacement, and what to check.
- Gemini 3.8 Flash TTS tops Hume's VoiceEQ board: run your own test
Hume's blog lists Gemini 3.8 Flash TTS atop its Real-World VoiceEQ board. Why a vendor-run board is only a lead, and how to run a blind A/B on your script.
- Omni Flash 1.1: GA in the Gemini API, Preview on Agent Platform
Gemini API release notes record gemini-omni-1.1-flash as GA on Aug 27, 2026; an Agent Platform page title still says Preview. What to check first.
Written by Sume