Check an AI take says your script: hypit align and unmatched words
hypit align pairs each script token with the transcript words of a generated take, and lists unmatched words. What it measures and what it does not do.

hypit_align checks a generated take against your script: for one script.svml segment it pairs every script token with the transcript words of the take, returns measured start_seconds, end_seconds and score, and lists unmatched[] for words the take did not say. No argument carries a time, and an unmatched word stays untimed.
The result is hypit.take/1. The surface is flag-gated: it is listed where the matching environment flag allows them (development auto-on, production opt-in), so check tools_list or the OpenAPI before you build on one.
What do I need before calling it?
The take must have its own probe and transcribe. Then send understanding_id (the reference's), take_understanding_id, transcript_job_id, script and segment.
curl -X POST https://api.sume.com/v1/hypit-understand/align \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: hypit-align-001" \
-d '{"understanding_id":"<reference probe id>",
"take_understanding_id":"<take probe id>",
"transcript_job_id":"<take transcribe id>",
"script":"<script.svml text>","segment":"<segment id>"}'What comes back?
Per the docs:
- Every script token with the measured span and score of the transcript words it paired with (an authored Korean word may pair with two heard words).
cues[]from||breaks and role changes.selections{}andmoments{}resolved to word boundaries.unmatched[]: script words the take did not say.take.jsonandscript.svmlas artifacts; the bundle liststakes[].
How do I read unmatched words?
An unmatched word is a signal to regenerate the take or edit the script, not to force a time. Because alignment is measured from the take's transcript, a mispronounced brand name will show as unmatched or low score, which is exactly the check you want before burning captions.
| Code | Meaning |
|---|---|
hypit_align_script_invalid | script.svml did not parse; message carries the hypit_svml_* code and line |
hypit_align_segment_unknown | segment is not a segment of the script |
hypit_transcript_not_found | Transcript job is not a completed transcribe of this take |
Where does this sit in a remake?
The docs lay out the lane for steps 2 to 5 of a reference remake: write project notes, produce takes with the existing generation tools, measure each take with a probe and transcribe, write the script as SVML, align, check, snapshot and bake. Align is the measurement step that stops you captioning words the voice never said.
Because the surface is flag-gated, run it where hypit_align appears in your tool list, and fall back to checking a take's transcript by hand with video inspect where it does not.
What does it not do?
Align does not generate the take, and it does not render a video. In the Sume Agent host only, sume-agent__hypit_compose compiles a composition.svml from aligned takes into a check or a billed MP4 bake; align itself is unbilled.
Sources
Related posts
More in Media tools
- Join voiceover takes into one gapless track with Timeline audio
Concatenate up to 20 Sume-hosted voice takes with Timeline audio, with no seam silence and no re-synthesis, and re-base video starts from the returned offsets.
- Crop a 16:9 video to a centered 9:16 strip: the crop fractions
For a centered 9:16 crop of a 16:9 video, send crop x 0.3418, y 0, width 0.3164, height 1 to Sume video-filter. 1:1 and 4:5 values are in the table.
- Cyber Monday ad in 9:16, 1:1 and 16:9 from one clip with Timeline
Resize one offer video to vertical, square and landscape with Timeline 1.0 output width and height, fit blur or cover, for $0.30 across three renders.
- Edit Video in Sume: which models appear and why H3 Max does not
Sume's Edit Video mode lists Auto, Kling 3.0, Wan 3.0 and MiniMax H3 and needs a reference video. H3 Max and Grok Imagine are missing; the API differs.
Written by Sume