hypit Understand order: one probe, then parallel batches
Sume's hypit Understand order: probe alone, then transcribe, boundaries and tiles in one batch, then notes. Which verbs wait on the transcript.

Run probe alone first to get the understanding id, then send transcribe, boundaries and the overview tile calls as one parallel batch. Only phrase-based tiles, word labels and the notes wait for the transcript. A second transcribe is never needed.
This is the order Sume's hypit Understand docs give, and it matters because each verb runs in its own container against a Modal Volume, so the source is not re-downloaded for the tenth grid. The docs mark this lane dest only: the routes and hypit_* tools are listed where SUME_COM_HYPIT_UNDERSTAND_ENABLED allows them (development auto-on, production opt-in), so check tools_list or the OpenAPI before you build on one.
What are the stages?
The docs describe four stages after the probe.
| Stage | Verbs | Waits on |
|---|---|---|
| 0 | probe alone | Nothing; mints the understanding id |
| 1 | transcribe + boundaries + overview tile | Probe |
| 2 | Dense tile and cut, chosen from boundaries and overview | Stage 1 boundaries |
| 3 | tile with around {phrase}, word-labeled close reads | Transcript |
| 4 | notes (ANALYSIS.md, TIMELINE.md, PROGRESS.md), then the bundle | Transcript |
What does each artifact look like?
Every image, clip and JSON is a durable media.sume.com artifact listed on the job's artifacts[], so a later step binds it by id. GET /v1/hypit-understand/:id on a probe id returns the bundle (hypit.understanding/1, with takes[]). cut returns an MP4 of up to 120 s of [start, end) or joined keep[] spans, with a mapping[] back to source time.
What are the notes for?
notes writes ANALYSIS.md, TIMELINE.md and PROGRESS.md (the reading of the reference), or FORMAT.md and TREATMENT.md (why the format works and the target treatment). The light lint requires evidence links to be media.sume.com artifacts, and FORMAT.md needs a ## Captions section.
curl -X POST https://api.sume.com/v1/hypit-understand/probe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: hypit-probe-001" \
-d '{"video_url":"https://media.sume.com/artifacts/artf_demo/ref.mp4"}'What mistakes does the order prevent?
Running every verb in series is the common one: each verb is its own job, so serial calls add each container's start and run time together. Re-running transcribe to get word labels is the other, since word-labeled tiles take the existing transcript_job_id instead.
Idempotency is also worth planning. The host tools derive the key from the verb and its arguments; on REST you send your own Idempotency-Key.
When is reference ingest enough?
If you want shots, OCR and audio facts in one call, use reference ingest. Use hypit when you want to read a reference in stages with scored words and labeled grids. For a quick probe or unlabeled stills, video inspect is the lighter route.
Sources
Related posts
More in Agents
- How an agent picks a video model over hosted MCP
An agent reads video-router_models, picks an id, then calls generate_video with an idempotency_key and waits with jobs_wait. Omit the model for sume/auto.
- Get a typed podcast clip plan from an Agent Completion
Send a transcript and an output_schema to POST /v1/agent/completions, and get back start and end times for each clip, ready for video-trim.
- Poll a Sume Action run in Python: branch on next_action, not status
A small Python loop for a Sume Scheduled run that sleeps on poll_status, retries on retry_later and stops on none, then fetches the result.
- Cron or API call: the trigger type of a Sume schedule is fixed
A Sume schedule's trigger_type is set at creation. An API-only schedule can never gain a cadence, and a cron one can never become API-only. What to pick first.
Written by Sume