Wedding toast transcript: video inspect with sentence segments

Get a text keepsake of a toast video: video-inspect with transcribe true returns text, word timings and sentence segments for $0.01 per audio minute.

5 min readSume
All posts

To get a written keepsake of a wedding toast, call POST /v1/video-inspect with transcribe: true on the recorded video, and add segmentation: { "mode": "sentence" }. The response carries the transcript text, words[] with timings and gapless sentence segments[]. The public rate is $0.01 per audio minute, so a 4-minute toast is about $0.04 plus the inspect's own Modal compute.

Source: Video inspect, read 2026-10-04.

Request

Send frames: false to skip the stills when you only want text. Everything that tunes the transcript is nested under transcribe.

curl -X POST https://api.sume.com/v1/video-inspect \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: toast-transcript-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/toast.mp4",
    "frames": false,
    "transcribe": true,
    "language_code": "en",
    "duration_seconds": 240,
    "segmentation": { "mode": "sentence" }
  }\

Fields that matter

Set duration_seconds as a hint to size the hold: if you omit it, Sume reserves one minute, and the maximum hint is 600 seconds. The charge never exceeds the hold.

Transcribe fields on video-inspect (read 2026-10-04)
FieldMeaningLimit
transcribeRuns Sume STT 1.0 on the clip's audioBoolean
language_codeSTT hint such as en or koOmit for auto-detect
duration_secondsBilling hint for the reservationOmitted reserves 1 minute; at most 600
segmentation.modesentence returns gapless sentence segmentssilence_split_seconds 0.2 to 3 is optional

Failure cases

If the video has no audio you get inspect_source_has_no_audio. Probe with frames: false and no transcript first, and read probe.has_audio. Any of language_code, segmentation or duration_seconds without transcribe: true is a 400 video_inspect_transcribe_required.

Proofread before printing

Speech-to-text will mishear names and in-jokes. Read the text against the video before you print it, and fix the names by hand. If you want the toast burned in as captions instead, POST /v1/video-captions takes a script_text that aligns to the speech.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume