Walmart video .vtt captions: Sume burns text in, you build the file
Walmart calls closed captions optional but recommended and wants a .vtt name. Sume burns captions into pixels and returns no .vtt; here is the transcript route.

Sume does not output a .vtt file. Walmart's rich media checklist, which it lists as last updated on Sep 17, 2026 (read 2026-10-08), says closed captioning is optional but recommended for accessibility, and that a closed-captioned file must have a filename ending in ".vtt". Sume's standalone caption job burns text into the video frames and returns a captioned video. You build the .vtt yourself from a transcript, which Sume can supply.
What Walmart asks for
The checklist is short. Video files must be .mp4, the size limit is 100MB, and captions are optional but recommended to meet accessibility guidelines and WCAG. It defines closed captions as time-synchronized text that reflects an audio track, so that viewers with a hearing disability can follow the video.
That definition matters. Text burned into the picture is always visible and cannot be switched off or read by a screen reader, so it is a different thing from a caption track. Burned-in text is still useful on a clip that plays muted, but it does not satisfy a request for a .vtt file.
What Sume returns
The caption endpoint, POST /v1/video-captions, takes a public HTTPS video_url and returns a captioned video. The docs say the API does not support SRT uploads, and they list no .vtt or .srt output. A caption job costs a flat $0.20 for videos up to 60 seconds, per the pricing note in the caption docs.
For a sidecar file you need timings. POST /v1/video-inspect with transcribe set to true runs Sume STT on the audio of a clip you already host on media.sume.com and returns a transcript with words, plus sentence segments when you set segmentation.mode to sentence. The public rate is $0.01 per audio minute, plus the job's Modal compute. A silent clip fails with inspect_source_has_no_audio, so check probe.has_audio first.
| Need | Walmart checklist | Sume surface | Price |
|---|---|---|---|
| Visible text on a muted clip | Not required | POST /v1/video-captions (burned in) | $0.20 per job, up to 60 s |
| Caption file for upload | Filename must end in .vtt | None. Build it from the transcript | You write the file |
| Timings for the file | Time-synchronized text | POST /v1/video-inspect, transcribe true | $0.01 per audio minute plus compute |
| Video container | Must be .mp4 | Confirm the delivered file is .mp4 before upload | None |
curl -X POST https://api.sume.com/v1/video-inspect \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: walmart-vtt-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"frames": false,
"transcribe": true,
"segmentation": { "mode": "sentence" }
}'A practical order of work
Render the clip, then run the inspect call with frames set to false so you pay for the transcript and not for stills. Turn each sentence segment into one cue in a WebVTT text file, name it so it ends in .vtt, and upload it beside the MP4 in Walmart's Media Library.
Read the transcript before you publish. Speech-to-text can mis-hear product names, and the caption file is the text that accessibility software reads. If you also want burned-in text for muted autoplay, send script_text to the caption job so the burned words follow your script and not the recognizer.
- Walmart's page does not say whether it reviews the .vtt text, so treat the file as part of the listing copy.
- A clip with music but no speech has nothing to transcribe. See the music-only caption rule.
- Keep the 100MB limit in mind. Walmart says it is typically but not always accepted.
What this does not cover
Walmart's page is the only source for the rules above, and the caption file format itself is outside it. Check the Media Library upload screen for how it pairs the .vtt with the video. Sume's docs describe no Walmart integration, so the upload is manual or through your own tooling.
Sources
Related posts
More in Use cases
- Wan 3.0 clip as a Genjutsu source: $13.80 for 30 s at 480p
Use a 30-second Wan 3.0 clip as the motion source for Higgsfield Genjutsu: the two-step cost at 480p and 720p, and the request for each step.
- Wedding save-the-date video with Omni: 8 seconds, vertical
A 9:16 save-the-date from an engagement photo with Gemini Omni Flash: $1.00 at 720p on Sume, with names and date added as captions, not generated text.
- Weekly TikTok trend sweep: 3 regions x 4 keywords is $1.20
Loop four keywords over US, GB and KR with Sume trending search: 12 calls, $1.20 a week. A short Python script and a month of cost.
- Which caption style wins? Test slam, punch, tiktok-green
Burn three caption styles onto one clip for $0.60 and compare them. Latin text works on all three; Korean text is rejected on these styles. Runnable loop.
Written by Sume