reference_ingest_stt_required: speech.language_code without STT
speech.language_code and duration_seconds only apply when the read transcribes. Set speech.allow_billed_stt true, use reference_remix, or drop them.

reference_ingest_stt_required is returned when a reference-ingest request sends speech.language_code or duration_seconds on a read that does not transcribe. The fix is one of two. Either make the read transcribe, with speech.allow_billed_stt: true or by leaving it out for purpose: reference_remix, or drop the two fields. The API hint is set_allow_billed_stt_or_drop_field, and the response names the offending field.
Reference ingest is dest first: the docs say Sume lists POST /v1/reference-ingest and the tool reference_ingest only where SUME_COM_REFERENCE_INGEST_ENABLED permits them, with development auto-on and production opt-in.
Why the default matters
Per the reference ingest docs, speech.allow_billed_stt defaults to true for purpose: reference_remix and false for the other three purposes, brief_format, face_swap and qa. A request that sets purpose: qa and a language_code therefore fails: the read will not transcribe, so the language hint has nothing to apply to.
| purpose | allow_billed_stt omitted | language_code or duration_seconds |
|---|---|---|
| reference_remix | true | accepted |
| brief_format | false | reference_ingest_stt_required |
| face_swap | false | reference_ingest_stt_required |
| qa | false | reference_ingest_stt_required |
| any purpose with allow_billed_stt true | true | accepted |
Two valid bodies
The first body transcribes a remix reference and gives a language hint. The second is a QA read with no speech fields, which is allowed. Both need an Idempotency-Key header, and the clip must be a media.sume.com file of at most 300 seconds.
{
"video_url": "https://media.sume.com/artifacts/artf_demo/ref.mp4",
"purpose": "reference_remix",
"speech": { "language_code": "en" }
}
{
"video_url": "https://media.sume.com/artifacts/artf_demo/ref.mp4",
"purpose": "qa"
}What allowing STT costs
A read that transcribes adds the speech-to-text rate to the reservation, per started minute of the hint, or one minute when there is no hint. The docs add that the read transcribes only when the track is not silent and voice activity detection finds speech, and otherwise the charge settles to zero. So a silent Short reference costs nothing for STT even with the flag on, and the response warns stt_skipped_silent or stt_skipped_no_speech.
Catching it before submit
The refusal comes from request validation, so it fires before any media is read. A guard in your own client is cheap: if purpose is not reference_remix and speech.allow_billed_stt is not true, strip speech.language_code and duration_seconds. Then the request can only fail for reasons that concern the file. Log the stripped fields so that a missing language hint on a transcribing read is not a silent surprise. When a read does transcribe and you omit the language hint, speech-to-text detects the language itself, which is fine for most Shorts in one language and worth overriding for mixed-language clips.
When not to enable it
If you only need cuts and on-screen text, leave STT off: pick a purpose other than reference_remix, or set allow_billed_stt: false, and do not send speech fields. The shots and text tracks are not billed as speech. Do not pass duration_seconds as a general length field; it is only the reservation hint for transcription, with a ceiling of 300.
Sources
Related posts
More in Media tools
- Reference ingest text_tracks is not caption coverage: 5 sampled frames
text_tracks come from OCR on deduplicated frame states, at most five frames, not every caption. To know if a Short is captioned, check frames yourself.
- Reframe 16:9 to 9:16: crop fractions for left, center, right
A 16:9 to 9:16 crop is a width of 0.3164 of the frame. Use x 0, 0.3418 or 0.6836 for a left, centered or right subject in Sume video filter.
- Restaurant menu video music: a 30-second brief and cost
What to ask Lyria for under a 30-second restaurant menu video: tempo, instruments, a clean ending and the $0.225 audio cost on Sume with a Timeline render.
- Review reel of six 360p Omni drafts: one timeline render for 10 cents
Join six 360p Omni drafts into one review file with Sume Timeline 1.0 for $0.10, with silence as the spine and contain-fit so nothing is cropped.
Written by Sume