Reference ingest purpose: qa or remix decides who transcribes
Reference ingest purpose defaults to reference_remix, which transcribes speech. brief_format, face_swap and qa do not. An explicit allow_billed_stt wins.

On Sume's reference ingest, purpose is one of reference_remix, brief_format, face_swap or qa, and omitting it means reference_remix. That default has one billing effect: a remix read transcribes the clip's speech unless you opt out, while the other three purposes do not. Set speech.allow_billed_stt to true or false and that value wins for every purpose.
A typical path this month is trend search first, then a read of the winner. Exploding Topics' September 21 list puts massage comb at +12,800% (Exploding Topics, read 2026-10-04); you find videos with trending search, import the one you may reuse into your workspace, then ingest it. Pick the purpose at that last step on what you will do with the words.
The default by purpose
The rule lives in referenceIngestAllowsBilledStt in packages/api-contract/src/reference-ingest.ts: an explicit boolean wins, otherwise purpose (default reference_remix) equals reference_remix. The reference ingest docs say purpose is stored and not interpreted; this transcription default is the one place it matters.
| purpose | allow_billed_stt omitted | allow_billed_stt true | allow_billed_stt false |
|---|---|---|---|
| reference_remix (default) | Yes | Yes | No |
| brief_format | No | Yes | No |
| face_swap | No | Yes | No |
| qa | No | Yes | No |
Billed only when speech exists
Allowing STT reserves the Sume STT 1.0 per-minute rate (the docs list sume/video-inspect-1.0#transcript), but the worker transcribes only when the track is not silent and voice-activity detection finds speech. Otherwise it settles to zero and the manifest carries a stt_skipped_silent, stt_skipped_no_speech or stt_skipped_no_audio_track warning. A clip with no audio therefore costs nothing for the transcript either way.
One 400 to know
speech.language_code and duration_seconds only apply when the read transcribes. Send either on a read that does not, such as purpose: "qa" with no override, and the request is refused with reference_ingest_stt_required. The fix is either to set speech.allow_billed_stt: true or to drop the field.
def transcribes(purpose=None, allow_billed_stt=None):
if isinstance(allow_billed_stt, bool):
return allow_billed_stt
return (purpose or "reference_remix") == "reference_remix"
cases = [
(None, None),
("brief_format", None),
("qa", None),
("qa", True),
("reference_remix", False),
]
for purpose, override in cases:
print(purpose, override, "->", transcribes(purpose, override))
Reference ingest is dest first with a production opt-in. Check tools_list for reference_ingest before you send a read, and confirm any per-minute number in GET /v1/catalog.
Sources
Related posts
More in Media tools
- Reference ingest shot record: cut types, palette, luma, motion
Each reference-ingest shot carries cut_out type (hard, gradual, end), palette, luma, contrast and a motion class. Turn them into shot length and pacing numbers.
- Reference ingest source block: rotation, vfr and aspect
Before you remake a reference, read the manifest source block: display_aspect, rotation, fps, vfr, codec and has_audio_track. What each field should change.
- Reference ingest uncertain[]: five kinds and what to do
The reference-ingest manifest lists five uncertain kinds, each with a suggested next step. Read the list, look again once per entry, and skip the rest.
- Remove a video background: Firefly vs Sume options
Adobe Firefly can now remove a video background across a whole clip. The Sume docs list a cutout tool for stills, not video, so plan a clean backdrop instead.
Written by Sume