ACE-Step 1.5 vocal-to-background music: what Sume does instead
ACE-Step 1.5 lists a vocal-to-background-music mode, covers and repainting under MIT. Sume's music jobs are prompt-only, so here is how to get a vocal-free bed.

If you want to turn a song with vocals into a background track, ACE-Step 1.5 lists that as a feature, and Sume does not. The model card names "vocal-to-background music conversion" next to text-to-music, covers, repainting and fine-tuning, under an MIT license. Sume's music jobs take a text prompt and an optional image, so on Sume you ask for a bed without vocals instead of converting an existing song.
What does ACE-Step 1.5 say it can do?
The card is short and specific. These are its claims, not measurements we ran.
| Item | Model card says |
|---|---|
| License | MIT |
| Memory | Runs locally with less than 4GB of VRAM |
| Speed | Full song in under 2 seconds on an A100, under 10 seconds on an RTX 3090 |
| Languages | 50+ |
| Length | Short loops up to 10-minute compositions |
| Features | Text-to-music, covers, repainting and editing, vocal-to-background music, fine-tuning |
| Training data | Licensed, royalty-free or no-copyright, and synthetic data |
What does Sume accept as input?
Music 1.0 takes prompt (1 to 5000 characters) and an optional public HTTPS image_url. There is no audio input, no seed and no duration field, and the Music Router uses the same body plus an optional model. So an existing song cannot be uploaded and converted.
How do you get a vocal-free bed on Sume?
Write the exclusion into the positive prompt. The docs say Lyria does not support negative prompting and that a non-empty negative_prompt returns HTTP 400. Close the brief with "Instrumental, no vocals.", and add "no spoken word" when the bed goes under narration. Then check the result by ear, since the docs treat the brief as direction, not a guaranteed setting.
Tempo, key and instruments help the match: the music page suggests a number for tempo (for example 72 BPM), a key and two to four instruments with texture.
Which should you use?
Use ACE-Step locally when the job is converting or repainting a track you already have the rights to and you have a GPU. Use a Sume Music Router job when you need a new bed with a Sume-hosted URL, billed at the fixed Music price, that can go into a Timeline soundtrack with loop, gain_db and duck_db. Check the rights to any song you feed a converter, whatever the converter's license says.
What does the speed claim change?
The card says a full song in under 2 seconds on an A100 and under 10 seconds on an RTX 3090. That matters when you want dozens of takes to audition. It does not tell you the takes are good, and the numbers are the model card's, not ours. On Sume, a music job is asynchronous: the docs say sync and subscribe modes wait up to 30 seconds, and a song may take longer, so poll the job instead of waiting on the request.
Either way, plan to listen. A bed that sounds right alone can fight a voice, which is what the duck_db field in the Timeline soundtrack is for.
What does the training-data line tell you?
The card says the training data comprised licensed data, royalty-free or no-copyright data and synthetic data, and that generated music is available for commercial purposes. That is the vendor's statement, and it is worth saving with the date you read it. It is not a guarantee about any single output, and it says nothing about the songs you might feed into a conversion mode.
Sources
Related posts
More in Models
- AI image model news, Oct 3, 2026: what API callers must do
Read from vendor docs Oct 3: xAI retires grok-imagine-image-quality Nov 2; OpenAI lists no October image entry; Google and BFL show no new image model.
- AI video audio by model: toggle, always on, or none on Sume
Seedance, Wan 3.0 and Kling 3 let you switch audio; MiniMax H3 and Gemini Omni Flash always make it; Grok Imagine has none. What each row does.
- AI video looks fake? Check Seedance 2.5 frames for skin and light
ByteDance says Seedance 2.5 tuned textures, skin, eyes, lighting and stray subtitles. How to pull stills from a clip on Sume and check each one.
- How to choose an AI video model: five questions before you pin one
Length, resolution, audio, start image and references decide the model, not the leaderboard. Five questions mapped to Sume's video catalog.
Written by Sume