Zillow Showcase video autoplays on mute: add cue captions
A Showcase video plays muted as the second carousel image, so a silent walkthrough needs on-screen words. Burn authored cues in with Sume captions.

Zillow Showcase video autoplays on mute as the second image in the listing carousel, so a walkthrough with no on-screen words says nothing until someone taps. Sume video captions can burn authored cues onto a silent clip without speech-to-text, for $0.20 on a video up to 60 seconds.
Zillow's Showcase Getting Started Guide: Editing listings (read 2026-10-03) describes the autoplay and the 120-second maximum, and says videos cannot include persistent agent or brokerage logos, watermarks or advertising. It does not say anything about captions, so check it before you pick a look.
As with any Showcase file, the photographer uploads through Aryeo, and agents choose from what is uploaded. Sume makes the captioned file; it does not upload it.
Why does a muted autoplay need cues instead of speech captions?
A house walkthrough is usually silent or has a music bed. Sume's speech captions need audible speech: a silent clip fails as caption_no_speech with next_action: use_overlay_captions. The fix the docs give is to pass cues (or segments), each with text, start and end in seconds. That path skips speech-to-text and burns exactly your copy at exactly those times.
script_text, words, cues and segments are mutually exclusive, so pick one per request. SRT uploads are not supported; pass phrase-level text as cues instead.
What does a cue request look like?
Send a public HTTPS video URL and the cues. Three beats cover a short walkthrough: the room, the one fact worth reading, and the close.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: showcase-cues-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/walkthrough.mp4",
"style": "slam",
"cues": [
{"text": "South-facing kitchen", "start": 1, "end": 5},
{"text": "Island seats four", "start": 6, "end": 10},
{"text": "Garden through the glass doors", "start": 12, "end": 17}
]
}'How do I keep the words readable and the rules intact?
Zillow bars persistent marks, so keep each cue short-lived and about the home, never a phone number or brokerage line held on screen. A cue that changes every few seconds reads as a caption; a fixed line reads like a plate.
Style and placement are yours to edit. A style is a set of design tokens, and the optional design field overrides one request, for example placement.anchor_ratio to move the line, or colors.active. If footage is busy, the earlier post on dimming a clip before captions covers the video filter route.
| Setting | Value | Why |
|---|---|---|
| cues | text + start + end in seconds | Skips speech-to-text, so a silent clip works. |
| style | slam, punch, tiktok-green or a Hangul identity | Omit it and the wording decides: slam for Latin text. |
| language | optional hint | Only tells speech-to-text what to expect; unused with cues. |
| price | $0.20 up to 60 seconds | Fixed estimate; confirm in GET /v1/catalog. |
What can go wrong?
Korean copy on slam, punch or tiktok-green is a 400 (caption_hangul_text_latin_style), because those faces have no Hangul glyphs. Pick a Hangul style for Korean cues.
The video URL must be a fetchable public HTTPS file; signed, private and non-HTTPS URLs are rejected. If you need a 120-second tour cut first, run video trim and caption the result. The captioned file arrives as a job result; see Jobs and results. To restyle without a new transcription, pass source_caption_id instead of video_url; billing is unchanged, a restyle is still a render.
Sources
Related posts
More in Use cases
- AI album cover generator: square art at 3000×3000
Generate square album art, then upscale: Apple recommends at least 3000×3000. On Sume, generate 2400×2400 and upscale it 1.25× to reach 3000×3000.
- AI avatar for online course videos: build and update lessons
Use an AI avatar as your online course instructor: one reusable avatar, a short talking video per section, captions, and one Timeline join per lesson.
- Talking avatar for PowerPoint presentations, slide by slide
Make a talking avatar presenter for PowerPoint: one Sume clip per slide, up to 60 seconds each, in 16:9 or 4:3 to match the slide, inserted as MP4.
- AI avatar for YouTube videos: Shorts and long-form
Use an AI avatar in YouTube videos: a 9:16 talking video of up to 60 seconds for a Short, or 16:9 segments joined into one long-form video.
Written by Sume