Check whether a reference video is silent before adding music
Reference ingest reports audio.silent at a -60 LUFS gate, speech presence and beats, so you know whether to keep, replace or add a soundtrack before a remix.

To check whether a video is silent before you add music, run it through POST /v1/reference-ingest and read audio.silent. The docs define it as true when integrated loudness is at or below −60 LUFS, or the true peak is −inf, so a clip with an audio track that carries no signal counts as silent.
The endpoint is dev-first: the Reference ingest page, read 2026-09-29, says production is opt-in, so confirm it is listed for your workspace.
What else does the audio block report?
| Field | Meaning |
|---|---|
silent | Integrated loudness at or below −60 LUFS, or true peak −inf |
| Speech presence | Silero voice-activity detection |
| Beats | Found with librosa when the track is music |
| Transcript | Only with speech.allow_billed_stt, and only if the track is not silent and speech is found |
What should I do when it is silent?
The docs say a silent source means: plan new music and discard the source audio, and never request a transcript on it. If you ask for one anyway, the transcript is skipped with stt_skipped_silent and settles to zero. Warnings stt_skipped_no_speech and stt_skipped_no_audio_track cover the other skip cases.
How do I put music under the result?
Build the final video with Timeline 1.0. A render takes one audio spine and ordered video[] slots. A silent video with a music bed can use audio.mode: "silence" for the spine plus a soundtrack, though duck_db needs a real spine. The next posts cover ducking music under a voice-over and rendering a silent video.
Is the check billed?
No. The manifest is unbilled. Only the optional transcript reserves the sume/video-inspect-1.0#transcript per-minute rate, and it settles to what actually ran.
How does the optional transcript work?
Set speech.allow_billed_stt and Sume transcribes only when the track is not silent and voice-activity detection finds speech; otherwise it settles to zero. speech.language_code is a hint and duration_seconds (at most 300) sizes the reservation, but both need allow_billed_stt, or the call fails with reference_ingest_stt_required. The transcript comes back with words and sentence segments.
Sources
Related posts
More in Developers
- Remove background from image in Node.js (JavaScript API)
Remove an image background from Node.js: call a background-removal API with fetch on your server, poll the job, and save the transparent PNG.
- Remove background from image in Python with an API
Remove an image background in Python: POST the image URL with requests, poll the job, then save the transparent PNG. A full script and the price.
- Replicate API rate limits: 600 creates a minute, then 429
Replicate's API allows 600 prediction creates and 3,000 other requests per minute. Low credit and no card tighten it; over the limit you get a 429.
- Runway API rate limit: usage tiers, concurrency and 429s
Runway's API has no requests-per-minute limit. Usage tiers cap concurrency per model, generations per 24 hours and monthly spend.
Written by Sume