Reference ingest transcript: allow_billed_stt and skipped STT
Reference ingest only bills a transcript when you opt in and the clip has speech. What allow_billed_stt reserves and what the stt_skipped warnings mean.

Reference ingest returns a transcript only when you set speech.allow_billed_stt: true, and then only if the track is not silent and voice activity detection finds speech. Otherwise the reserved amount settles to zero and you get a warning instead of a charge.
That makes it safe to turn on across a batch of references where you do not know which ones have a voiceover.
What does the manifest's audio block contain?
audio has a loudness gate (silent when integrated loudness is at or below -60 LUFS, or the true peak is -inf), Silero VAD speech presence, librosa beats when the track is music, and, only with the opt-in, a Sume STT 1.0 transcript with words and sentence segments.
The rule from the docs: audio.silent: true means plan new music and discard the source audio, and never request STT on it.
What does it cost?
The manifest itself is unbilled. With speech.allow_billed_stt, Sume reserves the sume/video-inspect-1.0#transcript rate per ceil(minute) of your duration_seconds hint (one minute when absent), then settles to what ran. The public STT rate on the video-inspect page is $0.01 per audio minute; confirm live in GET /v1/catalog.
| Clip | Warning | Transcript | Billed |
|---|---|---|---|
| Has speech | none | Words and segments | Per ceil minute |
| Music only | stt_skipped_no_speech | None | Settled to zero |
| Silent track | stt_skipped_silent | None | Settled to zero |
| No audio track | stt_skipped_no_audio_track | None | Settled to zero |
How do I request it?
speech.language_code and duration_seconds are only accepted with allow_billed_stt; sending them alone answers reference_ingest_stt_required. duration_seconds is capped at 300 because the clip is.
curl -X POST https://api.sume.com/v1/reference-ingest \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: ref-stt-001" \
-d '{"video_url":"https://media.sume.com/artifacts/artf_demo/ref.mp4",
"speech":{"allow_billed_stt":true,"language_code":"en"},
"duration_seconds":60}'How should I use this across a batch of references?
Set allow_billed_stt: true on every ingest and let the speech gate decide. Because a silent or music-only clip settles to zero, the worst case for a batch is the minutes of clips that actually contain speech. Pass a duration_seconds hint close to the real length, because the reservation is per ceil minute of the hint and one minute when you omit it.
Then branch on the warnings in your own code:
- No transcript and
stt_skipped_silent: plan new music; the source audio is not worth keeping. - No transcript and
stt_skipped_no_speech: the track is music; read beats fromaudioif you want to cut to them. - Transcript present: use the sentence segments to draft a script for the remake, and the words if you need exact timing.
- Something missing in the manifest, such as
reference_ingest_output_missing:<file>: treat that part as absent, not empty.
When should I use video inspect instead?
If you only want words and not shots or OCR, video inspect with transcribe: true is the direct route and takes clips up to 1800 s, with the transcript hint capped at 600 s. It errors with inspect_source_has_no_audio on a silent clip, where reference ingest warns and moves on.
Where does the cost show up in usage?
The reservation appears in the usage ledger like any other job. A job that skipped transcription settles to zero, and the warnings in the manifest tell you why. A scoped usage read by job id shows exactly what that one ingest cost.
If you run ingest as part of an agent workflow, the agent's own turns are separate rows in the same scope. The docs describe how the summary folds every row the scope caused, so the per-run figure includes the agent turns as well as the ingest.
For budgeting, assume zero for the manifest and cap the transcript at $0.05 per clip of up to 300 seconds.
Sources
Related posts
More in Pricing
- Runway Gen-4 Turbo and Wan3 Prime cost per second in dollars
At $0.01 per credit, Runway Gen-4 Turbo is $0.05 per second and Wan3 Prime is $0.068, $0.14 and $0.28 at 480p, 720p and 1080p. How to compare on Sume.
- Grok Imagine Video 1.5 on Runway: 10, 16, 29 credits a second
Runway bills grok_imagine_1_5 at 10, 16 or 29 credits per second by resolution, plus 1 per reference. Sume takes grok-imagine-video-1.5 as image-to-video.
- Runway real-time avatar cost per minute vs a rendered avatar video
Runway bills real-time avatars at 2 credits upfront plus 2 per 6 seconds, about $0.22 a minute. Sume's rendered avatar video is $0.184 to $0.55 per second.
- Seedance 2.5 token pricing on fal: formula and a 30-second 720p clip
fal bills Seedance 2.5 as height x width x seconds x 24 / 1024 tokens at $0.0214 per 1,000. Worked examples, a rounding gap on the page, and what Sume lists.
Written by Sume