hypit Understand order: one probe, then parallel batches

Sume's hypit Understand order: probe alone, then transcribe, boundaries and tiles in one batch, then notes. Which verbs wait on the transcript.

5 min readSume
All posts

Run probe alone first to get the understanding id, then send transcribe, boundaries and the overview tile calls as one parallel batch. Only phrase-based tiles, word labels and the notes wait for the transcript. A second transcribe is never needed.

This is the order Sume's hypit Understand docs give, and it matters because each verb runs in its own container against a Modal Volume, so the source is not re-downloaded for the tenth grid. The docs mark this lane dest only: the routes and hypit_* tools are listed where SUME_COM_HYPIT_UNDERSTAND_ENABLED allows them (development auto-on, production opt-in), so check tools_list or the OpenAPI before you build on one.

What are the stages?

The docs describe four stages after the probe.

hypit Understand stages (Sume docs, read 2026-10-02)
StageVerbsWaits on
0probe aloneNothing; mints the understanding id
1transcribe + boundaries + overview tileProbe
2Dense tile and cut, chosen from boundaries and overviewStage 1 boundaries
3tile with around {phrase}, word-labeled close readsTranscript
4notes (ANALYSIS.md, TIMELINE.md, PROGRESS.md), then the bundleTranscript

What does each artifact look like?

Every image, clip and JSON is a durable media.sume.com artifact listed on the job's artifacts[], so a later step binds it by id. GET /v1/hypit-understand/:id on a probe id returns the bundle (hypit.understanding/1, with takes[]). cut returns an MP4 of up to 120 s of [start, end) or joined keep[] spans, with a mapping[] back to source time.

What are the notes for?

notes writes ANALYSIS.md, TIMELINE.md and PROGRESS.md (the reading of the reference), or FORMAT.md and TREATMENT.md (why the format works and the target treatment). The light lint requires evidence links to be media.sume.com artifacts, and FORMAT.md needs a ## Captions section.

curl -X POST https://api.sume.com/v1/hypit-understand/probe \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: hypit-probe-001" \
  -d '{"video_url":"https://media.sume.com/artifacts/artf_demo/ref.mp4"}'

What mistakes does the order prevent?

Running every verb in series is the common one: each verb is its own job, so serial calls add each container's start and run time together. Re-running transcribe to get word labels is the other, since word-labeled tiles take the existing transcript_job_id instead.

Idempotency is also worth planning. The host tools derive the key from the verb and its arguments; on REST you send your own Idempotency-Key.

When is reference ingest enough?

If you want shots, OCR and audio facts in one call, use reference ingest. Use hypit when you want to read a reference in stages with scored words and labeled grids. For a quick probe or unlabeled stills, video inspect is the lighter route.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume