Word-level transcript with confidence scores: hypit transcribe
hypit transcribe returns each word with start, end and a 0-1 score using WhisperX, or Sume STT 1.0 without scores. Engines, checkpoints, billing and errors.

To get a word-timed transcript with a confidence score on every word, run transcribe in Sume's hypit Understand lane with the default engine: whisperx. Each word carries start_seconds, end_seconds and a score from 0 to 1; the alternative scribe_v2 engine is Sume STT 1.0, which has no per-word score.
hypit Understand is flag-gated, and its docs page is headed "Dest only": it is listed where the flag allows it (auto-on in development, opt-in in production), so a production workspace may not have it. Check tools_list or the OpenAPI to see whether your workspace can call it before you build on it.
What does the transcript look like?
The result is hypit.transcript/1: passages[], each with words[]. The WhisperX engine is Hypit's own upstream pipeline, faster-whisper ASR plus wav2vec2 alignment, running on the separate hypit-media Modal app. The score is useful as a filter: low-score words are where the aligner was unsure, so those are the ones to check before you caption or cut on them.
Which engine and checkpoint should I pick?
Language must be explicit. There is no model field on this route, because model is the public model id.
| Option | Values | Notes |
|---|---|---|
engine | whisperx (default) | Per-word score; unbilled |
engine | scribe_v2 | Sume STT 1.0, no per-word score; needs allow_billed_stt: true |
checkpoint | small (default) | int8 on a CPU container |
checkpoint | medium, large-v3-turbo, large-v3 | Opt-in, run on the GPU |
How do I call it?
Probe the clip once to get the understanding id (the probe's own job id), then transcribe against it.
curl -X POST https://api.sume.com/v1/hypit-understand/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: hypit-tr-001" \
-d '{"understanding_id":"<probe job id>","language":"en",
"engine":"whisperx","checkpoint":"small"}'How do I use the scores in practice?
The docs define the score as a 0 to 1 value on each aligned word and do not publish a recommended cutoff, so pick your own from a sample of your audio. A simple approach is to list every word under your cutoff and open a labeled tile at those times before trusting the transcript for captions or cuts.
Two things to remember: scores exist only on the WhisperX engine, and the small checkpoint is int8 on a CPU container, so a noisy or fast-talking reference is a reason to try large-v3-turbo or large-v3 on the GPU, which is opt-in.
What errors should I expect?
hypit_transcribe_checkpoint_conflict is a checkpoint sent with scribe_v2, since the checkpoint is a WhisperX knob. hypit_understanding_not_found means the id is not a completed probe of this workspace. source_too_long_for_hypit_understand is a clip over 300 s. With scribe_v2, billing is the STT per-minute rate on the probe's duration, settled to zero when the track is silent.
A second transcribe is never needed: later verbs bind to the transcript job id.
Sources
Related posts
More in Media tools
- Keyframe trim starts early: fix it with actual_start_seconds
A video-trim with precision keyframe can begin a GOP before your start time. The result reports actual_start_seconds; use it to re-base the next step.
- Korean caption styles compared: weight-shift to editorial-emphasis
Sume has six Hangul caption identities plus korean-ad. Compare black-outline, weight-shift, highlight, pill-karaoke, clip-wipe and editorial-emphasis.
- Korean subtitle fonts for a video API: 29 faces on Sume
Sume's caption font field names 29 Hangul faces, all SIL OFL 1.1, from Pretendard to Gmarket Sans. Which suit which style, and what fails on a Latin style.
- Add a launch date plate to a teaser video with compose overlay
Put a launch-date card on the bottom of a teaser with Timeline compose overlay: width_ratio, margin_ratio, a 300-second ceiling and $0.02 a job.
Written by Sume