How to make AI ASMR videos: prompt the picture and the sound

Use a video model that generates sound with the picture, and describe the close-up action and its sounds in the prompt. Models, prompts, and length.

5 min readSume
All posts

To make AI ASMR videos, use a video model that generates sound together with the picture, and write a prompt that describes both the close-up action (a knife slicing, fingers tapping, a bite into something crunchy) and the sound it should make. Then listen before you post: the sound is generated with the clip, and nothing guarantees it matches your prompt.

Facts come from Sume's Video generation, Video Router, Audio detach, and Timeline 1.0 docs and the catalog behind GET /v1/videos/models, read on 2026-09-28.

Which AI video generator makes ASMR sound?

On Sume, MiniMax H3, MiniMax H3 Max and Gemini Omni Flash 1.1 always generate sound; Seedance, Wan 3.0 and Kling 3 do unless you send generate_audio: false; and Grok Imagine Video 1.5 makes none. Which video models generate sound covers the generate_audio field and the values each model refuses.

From Video generation, Video Router, and the catalog behind GET /v1/videos/models, read 2026-09-28.
ModelSoundClip length
minimax-h3, minimax-h3-maxAlways, native stereo5–15 s
gemini-omni-flash-1.1Always, native synced3–10 s
seedance-2.5Optional, on by default4–30 s
seedance-2, seedance-2-fast, seedance-2-miniOptional, on by default4–15 s
wan-3.0Optional, on by default2–30 s
kling-3Optional, on by default4–15 s
grok-imagine-video-1.5None4–15 s

How do I make an AI ASMR video?

  • Pick a model with sound from the table, and send aspect_ratio: "9:16" for a vertical video.
  • Describe the picture: an extreme close-up or macro shot, a static camera, the material and its texture, and one slow action.
  • Describe the sound in the same prompt: what it should sound like, that the room is quiet, and no music or voice if you want only the object.
  • To show a specific object, send its photo as the first_frame in frame_images, at a public HTTPS URL.
curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: asmr-glass-kiwi-001" \
  -d '{
    "model": "gemini-omni-flash-1.1",
    "prompt": "Extreme close-up of a knife slowly slicing a glass kiwi on a wooden board. Static macro shot, soft studio light. Each cut makes a crisp glassy crunch; quiet room, no music, no voice",
    "aspect_ratio": "9:16",
    "resolution": "720p",
    "duration": 8
  }'

What should an AI ASMR prompt say?

Name the shot, the object, the action, and the sound, in that order. Three starting points to adapt:

  • Cutting: “Extreme close-up of a knife slicing a glass apple on a marble counter, slow even cuts, a crisp glassy crackle with each slice, quiet room, no music.”
  • Food: “Macro shot of a fork breaking into a crunchy honeycomb bar, honey dripping, a loud crisp crunch and a sticky pull, no music, no voice.”
  • Tapping: “Close-up of fingernails tapping slowly on a wooden box, then a glass jar, soft hollow taps and gentle clinks, static camera, warm light.”

How do I make a longer ASMR video without losing the sound?

Join several clips, and carry their sound yourself. In current code a Timeline 1.0 render takes sound only from its audio spine and an optional soundtrack, so each clip's own audio is dropped. Detach each clip's audio first with POST /v1/audio-detach, which returns a new WAV file by default along with its duration_seconds. Then list those files in order as audio.parts[], up to 20 slices joined without gaps, next to the matching clips in video[]. The parts must add up to at least audio.duration_seconds:

curl -X POST https://api.sume.com/v1/timeline-1.0/render \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: asmr-join-001" \
  -d '{
    "audio": {
      "duration_seconds": 16,
      "parts": [
        { "url": "https://media.sume.com/artifacts/artf_demo/cut-1.wav" },
        { "url": "https://media.sume.com/artifacts/artf_demo/cut-2.wav" }
      ]
    },
    "video": [
      { "source_url": "https://media.sume.com/artifacts/artf_demo/cut-1.mp4", "start": 0, "duration": 8 },
      { "source_url": "https://media.sume.com/artifacts/artf_demo/cut-2.mp4", "start": 8, "duration": 8 }
    ]
  }'

What are the limits and costs?

  • The docs make no promise about how closely the generated sound follows your prompt.
  • Detach and Timeline read only this workspace's media.sume.com files, such as the clips Sume generated for you.
  • One render runs 1–1,800 seconds, and the default output is 1080×1920.
  • Clips are billed per model at the provider's list price × 1.25, plus a 5.5% agent fee by default. On kling-3, sound raises the rate from $0.14 to $0.21 per second.
  • Each detach is $0.01 per job, and the render reserves $0.10 per output minute, rounded up to whole minutes; see API pricing.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume