Check an AI take says your script: hypit align and unmatched words

hypit align pairs each script token with the transcript words of a generated take, and lists unmatched words. What it measures and what it does not do.

5 min readSume
All posts

hypit_align checks a generated take against your script: for one script.svml segment it pairs every script token with the transcript words of the take, returns measured start_seconds, end_seconds and score, and lists unmatched[] for words the take did not say. No argument carries a time, and an unmatched word stays untimed.

The result is hypit.take/1. The surface is flag-gated: it is listed where the matching environment flag allows them (development auto-on, production opt-in), so check tools_list or the OpenAPI before you build on one.

What do I need before calling it?

The take must have its own probe and transcribe. Then send understanding_id (the reference's), take_understanding_id, transcript_job_id, script and segment.

curl -X POST https://api.sume.com/v1/hypit-understand/align \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: hypit-align-001" \
  -d '{"understanding_id":"<reference probe id>",
       "take_understanding_id":"<take probe id>",
       "transcript_job_id":"<take transcribe id>",
       "script":"<script.svml text>","segment":"<segment id>"}'

What comes back?

Per the docs:

  • Every script token with the measured span and score of the transcript words it paired with (an authored Korean word may pair with two heard words).
  • cues[] from || breaks and role changes.
  • selections{} and moments{} resolved to word boundaries.
  • unmatched[]: script words the take did not say.
  • take.json and script.svml as artifacts; the bundle lists takes[].

How do I read unmatched words?

An unmatched word is a signal to regenerate the take or edit the script, not to force a time. Because alignment is measured from the take's transcript, a mispronounced brand name will show as unmatched or low score, which is exactly the check you want before burning captions.

Align errors (Sume docs, read 2026-10-02)
CodeMeaning
hypit_align_script_invalidscript.svml did not parse; message carries the hypit_svml_* code and line
hypit_align_segment_unknownsegment is not a segment of the script
hypit_transcript_not_foundTranscript job is not a completed transcribe of this take

Where does this sit in a remake?

The docs lay out the lane for steps 2 to 5 of a reference remake: write project notes, produce takes with the existing generation tools, measure each take with a probe and transcribe, write the script as SVML, align, check, snapshot and bake. Align is the measurement step that stops you captioning words the voice never said.

Because the surface is flag-gated, run it where hypit_align appears in your tool list, and fall back to checking a take's transcript by hand with video inspect where it does not.

What does it not do?

Align does not generate the take, and it does not render a video. In the Sume Agent host only, sume-agent__hypit_compose compiles a composition.svml from aligned takes into a check or a billed MP4 bake; align itself is unbilled.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume