Can AI make a song from a voice memo or a video?
Suno v6 says it can start from a voice memo, image or video. Sume's music API takes a text prompt and one optional image, so audio and video stay outside it.

Yes in Suno, no in Sume. Suno's v6 post says you can start a song from text, audio, images or video, including a voice memo. Sume's music API takes a text prompt and one optional image; it has no field for an audio file or a video, so a voice memo has to be described in words instead.
Suno's wording is quoted from Introducing v6. Sume's fields come from the Music 1.0 and Music Router docs. Both were read 2026-09-29.
What does Suno say it can start from?
The post is dated Sep 9, 2026, and lists the inputs in one sentence with an example request.
| Suno's stated capability | Its own wording |
|---|---|
| Text, audio, images and video | Create with text, audio, images and video. Start with a written idea, voice memo, visual or video and turn it into music. |
What can Sume's music API take as input?
The docs give two inputs: a prompt of 1 to 5000 characters and an optional image_url, which must be a public HTTPS image. Nothing in either docs page lists an audio or video input. The docs call the image optional visual conditioning.
| Input | Sume music API |
|---|---|
| Written idea | prompt, 1 to 5000 characters |
| Image | image_url, optional, public HTTPS |
| Voice memo or other audio | Not listed in the docs |
| Video | Not listed in the docs |
How do I use a voice memo anyway?
Listen to it and write down what you hear as prompt text: the tempo, the key if you know it, the mood, the instruments and the words you sang. The Music 1.0 docs say the prompt steers the song, so the more of the memo you put into words, the closer the new track can get. It will be a new song that follows your description, not your recording turned into music.
- Write the lyrics you sang into the prompt under section labels such as
[Verse]and[Chorus]. - Say the length in the prompt.
durationandduration_secondsare rejected. - Read
result.lyricsafter the job, which the docs call model-reported, and listen before you rely on it.
Does the image field replace the audio input?
No. image_url is separate from anything you sing or hum, and the docs describe no way to send a recording next to it. If you want your own recorded voice in the final piece, you can lay it over a generated track afterwards; mixing a voice with background music shows how the timeline handles that.
How do I use a video?
Pick one still from the video, put it in image_url, and describe the pacing in the prompt. The music for a video post covers that route, and image to music covers the image field on its own.
Sources
Related posts
More in Media tools
- Mastodon video size limit: 99 MB, one video per post
Mastodon's docs allow a video of up to 99 MB (MP4, M4V, MOV or WebM), one per post, transcoded to H.264 at up to 1300 kbps. What that means for your export.
- Maximum video length for Sume's media tools: 300, 900 or 1800 seconds
Each Sume media tool has its own length cap: video filter and frames 300 s, trim and audio detach 1800 s in and 900 s out, compose 300 s. The full table.
- Microsoft Advertising video ads: specs for 1 to 3 stars
Microsoft Advertising rates Online Video ads 1–3 stars: 6–90 s at 2,500+ kbps for one star, exact lengths and 8,000–10,000+ kbps for more.
- Open Graph image size: 1200 × 630 and the 1.91:1 ratio
Meta recommends Open Graph images of at least 1200 × 630 px, close to 1.91:1 and under 8 MB, with 200 × 200 the minimum. How to make one per page.
Written by Sume