reference_ingest_stt_required: speech.language_code without STT

speech.language_code and duration_seconds only apply when the read transcribes. Set speech.allow_billed_stt true, use reference_remix, or drop them.

4 min readSume
All posts

reference_ingest_stt_required is returned when a reference-ingest request sends speech.language_code or duration_seconds on a read that does not transcribe. The fix is one of two. Either make the read transcribe, with speech.allow_billed_stt: true or by leaving it out for purpose: reference_remix, or drop the two fields. The API hint is set_allow_billed_stt_or_drop_field, and the response names the offending field.

Reference ingest is dest first: the docs say Sume lists POST /v1/reference-ingest and the tool reference_ingest only where SUME_COM_REFERENCE_INGEST_ENABLED permits them, with development auto-on and production opt-in.

Why the default matters

Per the reference ingest docs, speech.allow_billed_stt defaults to true for purpose: reference_remix and false for the other three purposes, brief_format, face_swap and qa. A request that sets purpose: qa and a language_code therefore fails: the read will not transcribe, so the language hint has nothing to apply to.

when STT hints are accepted (Sume docs and schema read 2026-10-05)
purposeallow_billed_stt omittedlanguage_code or duration_seconds
reference_remixtrueaccepted
brief_formatfalsereference_ingest_stt_required
face_swapfalsereference_ingest_stt_required
qafalsereference_ingest_stt_required
any purpose with allow_billed_stt truetrueaccepted

Two valid bodies

The first body transcribes a remix reference and gives a language hint. The second is a QA read with no speech fields, which is allowed. Both need an Idempotency-Key header, and the clip must be a media.sume.com file of at most 300 seconds.

{
  "video_url": "https://media.sume.com/artifacts/artf_demo/ref.mp4",
  "purpose": "reference_remix",
  "speech": { "language_code": "en" }
}

{
  "video_url": "https://media.sume.com/artifacts/artf_demo/ref.mp4",
  "purpose": "qa"
}

What allowing STT costs

A read that transcribes adds the speech-to-text rate to the reservation, per started minute of the hint, or one minute when there is no hint. The docs add that the read transcribes only when the track is not silent and voice activity detection finds speech, and otherwise the charge settles to zero. So a silent Short reference costs nothing for STT even with the flag on, and the response warns stt_skipped_silent or stt_skipped_no_speech.

Catching it before submit

The refusal comes from request validation, so it fires before any media is read. A guard in your own client is cheap: if purpose is not reference_remix and speech.allow_billed_stt is not true, strip speech.language_code and duration_seconds. Then the request can only fail for reasons that concern the file. Log the stripped fields so that a missing language hint on a transcribing read is not a silent surprise. When a read does transcribe and you omit the language hint, speech-to-text detects the language itself, which is fine for most Shorts in one language and worth overriding for mixed-language clips.

When not to enable it

If you only need cuts and on-screen text, leave STT off: pick a purpose other than reference_remix, or set allow_billed_stt: false, and do not send speech fields. The shots and text tracks are not billed as speech. Do not pass duration_seconds as a general length field; it is only the reservation hint for transcription, with a ceiling of 300.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume