Word-level transcript with confidence scores: hypit transcribe

hypit transcribe returns each word with start, end and a 0-1 score using WhisperX, or Sume STT 1.0 without scores. Engines, checkpoints, billing and errors.

5 min readSume
All posts

To get a word-timed transcript with a confidence score on every word, run transcribe in Sume's hypit Understand lane with the default engine: whisperx. Each word carries start_seconds, end_seconds and a score from 0 to 1; the alternative scribe_v2 engine is Sume STT 1.0, which has no per-word score.

hypit Understand is flag-gated, and its docs page is headed "Dest only": it is listed where the flag allows it (auto-on in development, opt-in in production), so a production workspace may not have it. Check tools_list or the OpenAPI to see whether your workspace can call it before you build on it.

What does the transcript look like?

The result is hypit.transcript/1: passages[], each with words[]. The WhisperX engine is Hypit's own upstream pipeline, faster-whisper ASR plus wav2vec2 alignment, running on the separate hypit-media Modal app. The score is useful as a filter: low-score words are where the aligner was unsure, so those are the ones to check before you caption or cut on them.

Which engine and checkpoint should I pick?

Language must be explicit. There is no model field on this route, because model is the public model id.

hypit transcribe options (Sume docs, read 2026-10-02)
OptionValuesNotes
enginewhisperx (default)Per-word score; unbilled
enginescribe_v2Sume STT 1.0, no per-word score; needs allow_billed_stt: true
checkpointsmall (default)int8 on a CPU container
checkpointmedium, large-v3-turbo, large-v3Opt-in, run on the GPU

How do I call it?

Probe the clip once to get the understanding id (the probe's own job id), then transcribe against it.

curl -X POST https://api.sume.com/v1/hypit-understand/transcribe \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: hypit-tr-001" \
  -d '{"understanding_id":"<probe job id>","language":"en",
       "engine":"whisperx","checkpoint":"small"}'

How do I use the scores in practice?

The docs define the score as a 0 to 1 value on each aligned word and do not publish a recommended cutoff, so pick your own from a sample of your audio. A simple approach is to list every word under your cutoff and open a labeled tile at those times before trusting the transcript for captions or cuts.

Two things to remember: scores exist only on the WhisperX engine, and the small checkpoint is int8 on a CPU container, so a noisy or fast-talking reference is a reason to try large-v3-turbo or large-v3 on the GPU, which is opt-in.

What errors should I expect?

hypit_transcribe_checkpoint_conflict is a checkpoint sent with scribe_v2, since the checkpoint is a WhisperX knob. hypit_understanding_not_found means the id is not a completed probe of this workspace. source_too_long_for_hypit_understand is a clip over 300 s. With scribe_v2, billing is the STT per-minute rate on the probe's duration, settled to zero when the track is silent.

A second transcribe is never needed: later verbs bind to the transcript job id.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume