Suno mashup of two songs: what Sume's music API can do
Sume's music models take a text prompt and one optional image, not audio, so no mashup of two songs. You can generate two tracks and join them end to end.

Sume cannot mash up two existing songs, because its music models accept no audio input. What it can do is generate two tracks from text prompts and join them end to end into one gapless file with Timeline audio. That is a sequence, not a layered blend.
Suno's claim is from its v6 announcement; Sume's facts are from Music 1.0 and Timeline audio, read 2026-09-30.
What does Suno v6 say about mashups?
Suno's post lists, among v6 features, "Build a mashup from multiple sources in one request". The page does not state limits or licensing terms, and this post makes no claim about either.
What does Sume's music request accept?
The Music 1.0 docs say it accepts a text prompt and optional image conditioning. The request fields are prompt (1 to 5000 characters), optional image_url (public HTTPS), negative_prompt (unsupported when non-empty), metadata, mode, webhook_url and wait_timeout_seconds. There is no audio field.
The Music Router adds an optional model; routable ids are sume/music-auto (default), lyria-3.5 and lyria-3-pro.
| Input | Accepted on Sume? |
|---|---|
| Text prompt | Yes, 1 to 5000 characters |
| One image | Yes, public HTTPS image_url |
| Audio of an existing song | No field |
| Several sources in one request | No |
How do I get the closest result on Sume?
Generate each track as its own job, then join them. Timeline audio with operation: "concat" takes parts[] of 1 to 20 ordered Sume-hosted audio files, joins them sample-exact with no silence at the seams, and returns one audio_url. All parts must already be media.sume.com audio in your workspace (import any outside file first with POST /v1/media-imports), and send an Idempotency-Key.
What does a join not do?
A concat plays one part after the other. It does not overlay vocals from one song on the instrumental of another, so it is not a mashup in the producer's sense. For one-song alternatives and edits, see Suno alternative: music API and editing part of a song on Sume.
Sources
Related posts
More in Use cases
- Text to speech add pause: Flux, HeyGen and Gemini markers vs Sume
Deepgram Flux, HeyGen and Gemini TTS each use their own pause marker. Sume TTS has none, so use punctuation, speed, or a declared silence in Timeline.
- TikTok 3-minute videos: cut a longer clip to 180 seconds via API
TikTok's page says all creators can post 3-minute videos, some 5 or 10. Sume video-trim writes a 180 s MP4 from a longer clip so any creator can post it.
- TikTok AI voiceover label: generic TTS vs a real person's voice
TikTok's 2026-H2 guidelines: generic text-to-speech narration needs no label, but AI audio that mimics a real person's voice does. How that maps to Sume voices.
- TikTok AI label policy: does anime or cartoon video need one?
TikTok counts anime and cartoons as AI content but says artistic styles need no disclosure. Realistic-looking people or scenes still do. The edge case.
Written by Sume