Walmart music-only video: add a [MUSIC PLAYING] caption with Sume cues

Walmart's best practice for a music-only video is a [MUSIC PLAYING] caption, removed after about 10 seconds. Burn it with Sume caption cues for a flat $0.20.

5 min readSume
All posts

For a Walmart Marketplace video that has music and no speech, Walmart's best practice is to insert a caption that says "[MUSIC PLAYING]" when the music starts and to remove it after about 10 seconds. With Sume you burn that text with POST /v1/video-captions and a cues array, which skips speech recognition. A standalone caption job costs a flat $0.20 for videos up to 60 seconds.

The rule in Walmart's words

The rule sits in the accessibility guidelines for videos on page 2 of Walmart's Rich media: Technical requirements PDF, which the PDF lists as last updated on Sep 17, 2026 (read 2026-10-08). It names two cases. A music-only video gets "[MUSIC PLAYING]" at the start of the music. A video with no music, audio or sound gets "[NO SPEECH]" at the start. In both cases the caption may come off after about 10 seconds.

Note the square brackets. They are part of the text Walmart quotes, so keep them in the cue.

Burning it with cues

Speech-to-text on a clip with no speech fails with caption_no_speech and next_action use_overlay_captions. Cues avoid that path. The caption docs say you can send cues, or segments, with text, start and end in seconds, that Sume then burns at those times without running speech-to-text. You can send only one of script_text, words, cues and segments.

Set start to 0 and end to 10 for a bed that starts with the clip. If the music enters later, start the cue where it enters and end it about 10 seconds after. The style field is optional. If you leave it out, the caption text picks the style, which is slam for Latin text.

Walmart music caption rule and the Sume fields that carry it (read 2026-10-08)
ItemValueSource
Caption text, music only[MUSIC PLAYING]Walmart PDF, accessibility guidelines
When it appearsWhen the music startsWalmart PDF
When it can come offAfter about 10 secondsWalmart PDF
Sume fieldcues: [{text, start, end}] in secondsSume video captions docs
Speech-to-textNot run when you send cuesSume video captions docs
Price per job$0.20, videos up to 60 sSume video captions docs
curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: walmart-music-cue-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/example/clean.mp4",
    "cues": [
      { "text": "[MUSIC PLAYING]", "start": 0, "end": 10 }
    ]
  }'

Where it fits in a batch

Walmart says nothing about how many clips need this, only that it is the best practice for a music-only video. If your catalog clips share one music bed, the same cue works on every file, so a batch is one caption job per clip: 40 clips cost 40 x $0.20 = $8.00.

The caption job needs a public HTTPS video_url that Sume can fetch. It rejects localhost, private-network and signed URLs, so use the media.sume.com artifact URL of the render.

  • Cue text is not checked against the audio. Make sure the clip really has no speech.
  • If the clip has speech, caption it normally with script_text or the transcript instead.
  • The caption is burned in, so it cannot be removed after the render. Render a separate version for any channel that does not want it.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume