Reference ingest purpose: qa or remix decides who transcribes

Reference ingest purpose defaults to reference_remix, which transcribes speech. brief_format, face_swap and qa do not. An explicit allow_billed_stt wins.

5 min readSume
All posts

On Sume's reference ingest, purpose is one of reference_remix, brief_format, face_swap or qa, and omitting it means reference_remix. That default has one billing effect: a remix read transcribes the clip's speech unless you opt out, while the other three purposes do not. Set speech.allow_billed_stt to true or false and that value wins for every purpose.

A typical path this month is trend search first, then a read of the winner. Exploding Topics' September 21 list puts massage comb at +12,800% (Exploding Topics, read 2026-10-04); you find videos with trending search, import the one you may reuse into your workspace, then ingest it. Pick the purpose at that last step on what you will do with the words.

The default by purpose

The rule lives in referenceIngestAllowsBilledStt in packages/api-contract/src/reference-ingest.ts: an explicit boolean wins, otherwise purpose (default reference_remix) equals reference_remix. The reference ingest docs say purpose is stored and not interpreted; this transcription default is the one place it matters.

Does the read transcribe? (from the contract rule, read 2026-10-04)
purposeallow_billed_stt omittedallow_billed_stt trueallow_billed_stt false
reference_remix (default)YesYesNo
brief_formatNoYesNo
face_swapNoYesNo
qaNoYesNo

Billed only when speech exists

Allowing STT reserves the Sume STT 1.0 per-minute rate (the docs list sume/video-inspect-1.0#transcript), but the worker transcribes only when the track is not silent and voice-activity detection finds speech. Otherwise it settles to zero and the manifest carries a stt_skipped_silent, stt_skipped_no_speech or stt_skipped_no_audio_track warning. A clip with no audio therefore costs nothing for the transcript either way.

One 400 to know

speech.language_code and duration_seconds only apply when the read transcribes. Send either on a read that does not, such as purpose: "qa" with no override, and the request is refused with reference_ingest_stt_required. The fix is either to set speech.allow_billed_stt: true or to drop the field.

def transcribes(purpose=None, allow_billed_stt=None):
    if isinstance(allow_billed_stt, bool):
        return allow_billed_stt
    return (purpose or "reference_remix") == "reference_remix"

cases = [
    (None, None),
    ("brief_format", None),
    ("qa", None),
    ("qa", True),
    ("reference_remix", False),
]
for purpose, override in cases:
    print(purpose, override, "->", transcribes(purpose, override))

Reference ingest is dest first with a production opt-in. Check tools_list for reference_ingest before you send a read, and confirm any per-minute number in GET /v1/catalog.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume