Get a transcript of an AI video before TikTok Shop review
Run Sume video inspect with transcribe: true to read what an AI video says, at $0.01 per audio minute plus compute, before you post a TikTok Shop clip.

To get a transcript of an AI video before you post it, call POST /v1/video-inspect with transcribe: true on a media.sume.com clip. The transcript is billed at Sume's public STT rate of $0.01 per audio minute, on top of the inspect's compute, and it gives you the exact words so you can compare them to your listing.
That comparison matters for TikTok Shop. The US page bans exaggerated product effects and fabricated expert personas, and the UK page bans fabricating product capabilities. A voice-over that a model improvised can say something your script never did.
How to run it
Video inspect reads one clip that the workspace already owns on media.sume.com; it will not fetch from the open internet, so import first with POST /v1/media-imports. The default mode is sync, which waits up to 30 seconds for a 200 and otherwise returns 202, and you poll the job.
Send frames: false if you want no stills. Send duration_seconds (up to 600) as a hint so the reserve matches the clip; without it Sume reserves one minute. language_code is an STT hint, and with segmentation.mode: "sentence" you also get sentence segments.
curl -X POST https://api.sume.com/v1/video-inspect \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: inspect-transcript-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"frames": false,
"transcribe": true,
"language_code": "en",
"segmentation": {"mode": "sentence"}
}'What to do with the text
Read the transcript against the claims on your product page. Strike any line that promises a result, a medical benefit, or a comparison you cannot support. Then fix the clip: re-generate the voice, or trim the line out with Audio detach and Timeline if you only need to drop a sentence.
Keep the transcript with the post record. Both Shop pages ask for disclosure when AI plays a significant role in the content, and a stored transcript shows what the final video actually said.
| Item | Value |
|---|---|
| Source | One media.sume.com clip, up to 1800 s |
| STT rate | $0.01 per audio minute (confirm in GET /v1/catalog) |
Reserve without duration_seconds | 1 minute |
duration_seconds hint | Up to 600 s |
| No audio track | Fails with inspect_source_has_no_audio |
Edge cases
A silent clip fails with inspect_source_has_no_audio. If you send language_code, segmentation, or duration_seconds without transcribe: true, you get 400 video_inspect_transcribe_required. Inspect also bills its own container seconds, so the total is more than the STT line; the docs do not give a fixed number for that part.
Sources
Related posts
More in Developers
- Get transcript text from a captioned video: caption jobs return none
A Sume caption job returns the burned video, not the transcript. For text and word times run STT or video inspect with transcribe. Prices and a recipe.
- Go: cancel the rest of a batch after the first 402 on Sume
A Go fan-out that stops launching Wan 3.0 submits the moment one returns 402 insufficient_credits, using context.WithCancel and a 2-slot semaphore.
- Go: a context timeout stops waiting on a Sume job, not the job
context.WithTimeout cancels your status read, never the generation. A 29-line Go loop shows the DeadlineExceeded branch, and why the job id must be stored.
- Go client for a Sume bulk queue: transient polls and exit status
Create a Sume bulk queue from Go, poll the status_url with a doubling gap, treat 429 and 503 as transient, and exit 1 on failed items. Standard library only.
Written by Sume