How to make AI ASMR videos: prompt the picture and the sound
Use a video model that generates sound with the picture, and describe the close-up action and its sounds in the prompt. Models, prompts, and length.

To make AI ASMR videos, use a video model that generates sound together with the picture, and write a prompt that describes both the close-up action (a knife slicing, fingers tapping, a bite into something crunchy) and the sound it should make. Then listen before you post: the sound is generated with the clip, and nothing guarantees it matches your prompt.
Facts come from Sume's Video generation, Video Router, Audio detach, and Timeline 1.0 docs and the catalog behind GET /v1/videos/models, read on 2026-09-28.
Which AI video generator makes ASMR sound?
On Sume, MiniMax H3, MiniMax H3 Max and Gemini Omni Flash 1.1 always generate sound; Seedance, Wan 3.0 and Kling 3 do unless you send generate_audio: false; and Grok Imagine Video 1.5 makes none. Which video models generate sound covers the generate_audio field and the values each model refuses.
| Model | Sound | Clip length |
|---|---|---|
minimax-h3, minimax-h3-max | Always, native stereo | 5–15 s |
gemini-omni-flash-1.1 | Always, native synced | 3–10 s |
seedance-2.5 | Optional, on by default | 4–30 s |
seedance-2, seedance-2-fast, seedance-2-mini | Optional, on by default | 4–15 s |
wan-3.0 | Optional, on by default | 2–30 s |
kling-3 | Optional, on by default | 4–15 s |
grok-imagine-video-1.5 | None | 4–15 s |
How do I make an AI ASMR video?
- Pick a model with sound from the table, and send
aspect_ratio: "9:16"for a vertical video. - Describe the picture: an extreme close-up or macro shot, a static camera, the material and its texture, and one slow action.
- Describe the sound in the same prompt: what it should sound like, that the room is quiet, and no music or voice if you want only the object.
- To show a specific object, send its photo as the
first_frameinframe_images, at a public HTTPS URL.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: asmr-glass-kiwi-001" \
-d '{
"model": "gemini-omni-flash-1.1",
"prompt": "Extreme close-up of a knife slowly slicing a glass kiwi on a wooden board. Static macro shot, soft studio light. Each cut makes a crisp glassy crunch; quiet room, no music, no voice",
"aspect_ratio": "9:16",
"resolution": "720p",
"duration": 8
}'What should an AI ASMR prompt say?
Name the shot, the object, the action, and the sound, in that order. Three starting points to adapt:
- Cutting: “Extreme close-up of a knife slicing a glass apple on a marble counter, slow even cuts, a crisp glassy crackle with each slice, quiet room, no music.”
- Food: “Macro shot of a fork breaking into a crunchy honeycomb bar, honey dripping, a loud crisp crunch and a sticky pull, no music, no voice.”
- Tapping: “Close-up of fingernails tapping slowly on a wooden box, then a glass jar, soft hollow taps and gentle clinks, static camera, warm light.”
How do I make a longer ASMR video without losing the sound?
Join several clips, and carry their sound yourself. In current code a Timeline 1.0 render takes sound only from its audio spine and an optional soundtrack, so each clip's own audio is dropped. Detach each clip's audio first with POST /v1/audio-detach, which returns a new WAV file by default along with its duration_seconds. Then list those files in order as audio.parts[], up to 20 slices joined without gaps, next to the matching clips in video[]. The parts must add up to at least audio.duration_seconds:
curl -X POST https://api.sume.com/v1/timeline-1.0/render \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: asmr-join-001" \
-d '{
"audio": {
"duration_seconds": 16,
"parts": [
{ "url": "https://media.sume.com/artifacts/artf_demo/cut-1.wav" },
{ "url": "https://media.sume.com/artifacts/artf_demo/cut-2.wav" }
]
},
"video": [
{ "source_url": "https://media.sume.com/artifacts/artf_demo/cut-1.mp4", "start": 0, "duration": 8 },
{ "source_url": "https://media.sume.com/artifacts/artf_demo/cut-2.mp4", "start": 8, "duration": 8 }
]
}'What are the limits and costs?
- The docs make no promise about how closely the generated sound follows your prompt.
- Detach and Timeline read only this workspace's
media.sume.comfiles, such as the clips Sume generated for you. - One render runs 1–1,800 seconds, and the default output is 1080×1920.
- Clips are billed per model at the provider's list price × 1.25, plus a 5.5% agent fee by default. On
kling-3, sound raises the rate from $0.14 to $0.21 per second. - Each detach is $0.01 per job, and the render reserves $0.10 per output minute, rounded up to whole minutes; see API pricing.
Sources
Related posts
More in Use cases
- How to make AI history videos from a researched script
Make AI history videos: research the script, make a still or use an archival photo per scene, animate each, narrate, and label AI scenes.
- How to make an AI documentary video, chapter by chapter
Make an AI documentary video from a researched script: voice it in chapters, make a shot per beat, and assemble up to 30 minutes in one render.
- How to make an AI voiceover: from script to audio file
Make an AI voiceover in five steps: write the script, pick a voice, generate the speech, check its length, and download the file. The Sume API way.
- How to make an unboxing video with AI, step by step
An unboxing video shows hands opening the package and revealing the product. How to make one with AI from a box photo and a product photo.
Written by Sume