Wedding toast transcript: video inspect with sentence segments
Get a text keepsake of a toast video: video-inspect with transcribe true returns text, word timings and sentence segments for $0.01 per audio minute.

To get a written keepsake of a wedding toast, call POST /v1/video-inspect with transcribe: true on the recorded video, and add segmentation: { "mode": "sentence" }. The response carries the transcript text, words[] with timings and gapless sentence segments[]. The public rate is $0.01 per audio minute, so a 4-minute toast is about $0.04 plus the inspect's own Modal compute.
Source: Video inspect, read 2026-10-04.
Request
Send frames: false to skip the stills when you only want text. Everything that tunes the transcript is nested under transcribe.
curl -X POST https://api.sume.com/v1/video-inspect \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: toast-transcript-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/toast.mp4",
"frames": false,
"transcribe": true,
"language_code": "en",
"duration_seconds": 240,
"segmentation": { "mode": "sentence" }
}\Fields that matter
Set duration_seconds as a hint to size the hold: if you omit it, Sume reserves one minute, and the maximum hint is 600 seconds. The charge never exceeds the hold.
| Field | Meaning | Limit |
|---|---|---|
| transcribe | Runs Sume STT 1.0 on the clip's audio | Boolean |
| language_code | STT hint such as en or ko | Omit for auto-detect |
| duration_seconds | Billing hint for the reservation | Omitted reserves 1 minute; at most 600 |
| segmentation.mode | sentence returns gapless sentence segments | silence_split_seconds 0.2 to 3 is optional |
Failure cases
If the video has no audio you get inspect_source_has_no_audio. Probe with frames: false and no transcript first, and read probe.has_audio. Any of language_code, segmentation or duration_seconds without transcribe: true is a 400 video_inspect_transcribe_required.
Proofread before printing
Speech-to-text will mishear names and in-jokes. Read the text against the video before you print it, and fix the names by hand. If you want the toast burned in as captions instead, POST /v1/video-captions takes a script_text that aligns to the speech.
Sources
Related posts
More in Use cases
- Weekly toolbox talk as a 45-second avatar clip: 52 weeks priced
A weekly 45-second safety briefing from a Sume avatar on a worksite photo, captions on: 52 clips cost $440.96 on Standard. Tiers and script rules.
- Meta AI disclosure: which Sume outputs count, video and audio
Meta asks for disclosure of photorealistic video and realistic audio that was digitally created or altered. A sorting guide for video, avatar and music.
- Do YouTube Shorts AI tools label automatically? What the page says
YouTube says its own Shorts generative AI tools disclose automatically. Regional limits per tool, and what an API pipeline still has to disclose itself.
- Xiaohongshu note video: 1:1 or 9:16, 500 MB, 720p
A third-party guide to Xiaohongshu note video specs lists 9:16 or 1:1, up to 15 minutes, up to 500 MB and 720p minimum.
Written by Sume