AI backing track generator: an instrumental to sing or play over
Generate an AI backing track by prompt: tempo, key, form, no vocals. Sume returns one mixed MP3, no stems or click track, so trim and check the take by ear.

To generate an AI backing track, ask the Music Router for an instrumental in the key and tempo you sing or play in, say the song form in the prompt, and close with "Instrumental, no vocals." Sume returns one mixed audio file, normally an MP3 on media.sume.com. It does not return stems, a click track or a chord chart, and it has no duration or BPM field, so you steer both in the prompt and check them by ear.
That is enough for rehearsal, demos and a sing-along video. It is not enough if you need a drum-free version, an isolated bass, or a loop with an exact bar count. The sections below say what you can control, what you cannot, and how to cut the file down to the part you need.
What can the prompt control?
The Music docs list the levers: tempo as a number, key and mode, two to four instruments with texture, and an arc with one named moment. They also support section markers such as [0:00-0:30] Intro: ... inside the prompt. Map those to the form of your song and you have a backing track brief. The tempo and key below are examples, and the docs call every axis a creative direction rather than a guaranteed setting.
| You need | What Sume gives | Workaround |
|---|---|---|
| A fixed tempo | A BPM written in the prompt, not a parameter | Name the number, then tap-check the take |
| A fixed key | Key and mode as a prompt axis | Say "D major" and test with your instrument |
| A set length | No duration field; a duration is rejected | Say "a 2-minute track" and trim with Timeline audio split |
| Verse, chorus, bridge | Section markers inside the prompt | Write [0:00-0:20] Verse: ... lines |
| No lead vocal | An exclusion written into the positive prompt | Close with "Instrumental, no vocals"; negative_prompt returns 400 |
| Separate drums or bass | One mixed file only; no stems | Generate a second take without that instrument |
How do I write a brief that leaves room for a singer?
Here is a brief for a singer who wants a two-verse acoustic backing in G major. Notice that the lead line is left out, so there is room for the voice, and the form is spelled out section by section.
curl -X POST https://api.sume.com/v1/music-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: backing-g-major-001" \
-d '{
"prompt": "Backing track for a singer, 84 BPM, G major. Fingerpicked acoustic guitar, upright bass, soft brushed drums, no lead melody so the vocal has room. [0:00-0:08] Intro: guitar alone. [0:08-0:40] Verse: bass and brushes enter. [0:40-1:10] Chorus: fuller, strummed. [1:10-1:40] Verse two: back to fingerpicking. A 2-minute track. Instrumental, no vocals."
}'How do I check the take?
Check three things against your own playing. Tap the pulse against a metronome app for one section. Play a held G against the track, and listen for beating on the first chord. Then count the form against your song, because the engine may place a section a few seconds off. If it is wrong, change one line and run again. With no seed, each submit is a new take, and each accepted generation costs the fixed Music price.
Remember that no stems exist. If a backing track still has a guitar line that clashes with your part, regenerate with a different instrument list rather than hoping to mute it. Stability's Stable Audio 3.0 is also instrumental-only per its vendor page, but it is a separate product and not a Sume route.
How do I trim the intro or keep one section?
Timeline audio can slice a file into up to 20 ranges and return each as its own durable file, so you can drop the intro or keep only the chorus for practice. Each range has a start and an optional end; the output is wav by default and the job is a flat $0.01. The input must already be a media.sume.com audio file, which a Music Router artifact is.
curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: backing-split-001" \
-d '{
"operation": "split",
"url": "https://media.sume.com/artifacts/artf_demo/backing.mp3",
"ranges": [{"start": 8, "end": 40}, {"start": 40, "end": 70}]
}'How many takes should I make?
Because every submit is a fresh take, the cheap habit is to run three variants of one brief and pick by ear. The Music 1.0 page lists the fixed price as $0.125 per accepted generation and says it does not vary by prompt length, and the router docs say every router model charges that same fixed Music price, so three takes cost three times that. Give each take its own Idempotency-Key, since reusing a key for the same payload returns the original job instead of a new one.
Vary the one thing you are unsure of: the tempo number by a few BPM, the lead-free line, or the instrument list. Name each take in the metadata field, which Sume stores on the job and does not send to the provider, so you can tell which prompt produced which file when you listen back a week later.
What does Sume not do for a backing track?
Sume has no way to feed your own recording in, so a track cannot follow what you play. It also cannot make a click track or a chord chart, and it cannot return a version without the drums. If you need any of those, a DAW or a notation tool is the right place to start, and Sume is the quick route to a mood and a feel.
Google's Lyria page says tracks carry a SynthID watermark. If you post a performance over the track, say that the backing was AI-generated where the platform asks.
Sources
Related posts
More in Use cases
- AI character series on Shorts: avoid the same situation each time
YouTube's inauthentic content policy flags characters in identical situations with the same outcomes. How to keep an AI character and vary the story in Sume.
- AI classroom background music for lesson videos, under narration
Make calm instrumental music for a lesson video: a prompt that keeps vocals out, a Python script, and a Timeline bed that ducks under the teacher's voice.
- AI person in a Meta ad: is the label next to Sponsored?
Meta puts AI info next to Sponsored when its own tools make a photorealistic human. For outside tools it describes About this ad. How to add your own cue.
- AI image slideshow Shorts on YouTube: what narrative to add
YouTube's inauthentic content policy lists image slideshows with minimal narrative as not allowed. Turn stills into a Short with a real story, using Sume.
Written by Sume