AI backing track generator: an instrumental to sing or play over

Generate an AI backing track by prompt: tempo, key, form, no vocals. Sume returns one mixed MP3, no stems or click track, so trim and check the take by ear.

5 min readSume
All posts

To generate an AI backing track, ask the Music Router for an instrumental in the key and tempo you sing or play in, say the song form in the prompt, and close with "Instrumental, no vocals." Sume returns one mixed audio file, normally an MP3 on media.sume.com. It does not return stems, a click track or a chord chart, and it has no duration or BPM field, so you steer both in the prompt and check them by ear.

That is enough for rehearsal, demos and a sing-along video. It is not enough if you need a drum-free version, an isolated bass, or a loop with an exact bar count. The sections below say what you can control, what you cannot, and how to cut the file down to the part you need.

What can the prompt control?

The Music docs list the levers: tempo as a number, key and mode, two to four instruments with texture, and an arc with one named moment. They also support section markers such as [0:00-0:30] Intro: ... inside the prompt. Map those to the form of your song and you have a backing track brief. The tempo and key below are examples, and the docs call every axis a creative direction rather than a guaranteed setting.

Backing-track needs against what the Sume docs offer, read 2026-10-03
You needWhat Sume givesWorkaround
A fixed tempoA BPM written in the prompt, not a parameterName the number, then tap-check the take
A fixed keyKey and mode as a prompt axisSay "D major" and test with your instrument
A set lengthNo duration field; a duration is rejectedSay "a 2-minute track" and trim with Timeline audio split
Verse, chorus, bridgeSection markers inside the promptWrite [0:00-0:20] Verse: ... lines
No lead vocalAn exclusion written into the positive promptClose with "Instrumental, no vocals"; negative_prompt returns 400
Separate drums or bassOne mixed file only; no stemsGenerate a second take without that instrument

How do I write a brief that leaves room for a singer?

Here is a brief for a singer who wants a two-verse acoustic backing in G major. Notice that the lead line is left out, so there is room for the voice, and the form is spelled out section by section.

curl -X POST https://api.sume.com/v1/music-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: backing-g-major-001" \
  -d '{
    "prompt": "Backing track for a singer, 84 BPM, G major. Fingerpicked acoustic guitar, upright bass, soft brushed drums, no lead melody so the vocal has room. [0:00-0:08] Intro: guitar alone. [0:08-0:40] Verse: bass and brushes enter. [0:40-1:10] Chorus: fuller, strummed. [1:10-1:40] Verse two: back to fingerpicking. A 2-minute track. Instrumental, no vocals."
  }'

How do I check the take?

Check three things against your own playing. Tap the pulse against a metronome app for one section. Play a held G against the track, and listen for beating on the first chord. Then count the form against your song, because the engine may place a section a few seconds off. If it is wrong, change one line and run again. With no seed, each submit is a new take, and each accepted generation costs the fixed Music price.

Remember that no stems exist. If a backing track still has a guitar line that clashes with your part, regenerate with a different instrument list rather than hoping to mute it. Stability's Stable Audio 3.0 is also instrumental-only per its vendor page, but it is a separate product and not a Sume route.

How do I trim the intro or keep one section?

Timeline audio can slice a file into up to 20 ranges and return each as its own durable file, so you can drop the intro or keep only the chorus for practice. Each range has a start and an optional end; the output is wav by default and the job is a flat $0.01. The input must already be a media.sume.com audio file, which a Music Router artifact is.

curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: backing-split-001" \
  -d '{
    "operation": "split",
    "url": "https://media.sume.com/artifacts/artf_demo/backing.mp3",
    "ranges": [{"start": 8, "end": 40}, {"start": 40, "end": 70}]
  }'

How many takes should I make?

Because every submit is a fresh take, the cheap habit is to run three variants of one brief and pick by ear. The Music 1.0 page lists the fixed price as $0.125 per accepted generation and says it does not vary by prompt length, and the router docs say every router model charges that same fixed Music price, so three takes cost three times that. Give each take its own Idempotency-Key, since reusing a key for the same payload returns the original job instead of a new one.

Vary the one thing you are unsure of: the tempo number by a few BPM, the lead-free line, or the instrument list. Name each take in the metadata field, which Sume stores on the job and does not send to the provider, so you can tell which prompt produced which file when you listen back a week later.

What does Sume not do for a backing track?

Sume has no way to feed your own recording in, so a track cannot follow what you play. It also cannot make a click track or a chord chart, and it cannot return a version without the drums. If you need any of those, a DAW or a notation tool is the right place to start, and Sume is the quick route to a mood and a feel.

Google's Lyria page says tracks carry a SynthID watermark. If you post a performance over the track, say that the backing was AI-generated where the platform asks.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume