Remove filler words from a talking video with an API
Sume has no one-call filler remover. Transcribe with video inspect for word timings, then cut the clean ranges with video trim at $0.02 per job.

Sume has no single call that strips filler words. You can build it from documented parts: POST /v1/video-inspect with transcribe: true returns word timings, and POST /v1/video-trim cuts each range you keep into a new MP4. You decide which words are filler, and you join the pieces yourself.
What does HeyGen's Speech Cleanup do?
HeyGen's June 2026 notes say it "removes every filler word, awkward pause, and false start, then stitches the remaining footage into a single seamless take with no visible jump cuts." That automated stitching is HeyGen's feature; Sume's docs do not describe an equivalent.
How do I get word timings?
Send the clip, hosted on media.sume.com, to video inspect with transcribe: true. The result carries transcript with text, words[] and optional sentence segments[]. Transcription is billed at $0.01 per audio minute; probe and stills are unbilled. A clip with no audio track fails as inspect_source_has_no_audio. Details: Video inspect.
curl -X POST https://api.sume.com/v1/video-inspect \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: filler-inspect-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"frames": false,
"transcribe": true,
"duration_seconds": 120
}'How do I cut the clean ranges?
Scan words[] for the words you count as filler, and take the gaps between them as keep ranges. Send each keep range to video trim with start and exactly one of end or duration. The default precision is exact, a frame-accurate re-encode, which suits speech cuts. Then place the returned MP4s in order on a Timeline.
| Step | Route | Documented price |
|---|---|---|
| Word timings | POST /v1/video-inspect | $0.01 per audio minute |
| Cut one range | POST /v1/video-trim | $0.02 per job |
| Join the pieces | Timeline 1.0 | See the Timeline docs |
What are the limits?
Source video must be at most 1800 seconds for both routes, a trim output is 0.2-900 seconds, and each trim is one job, so a talk with 40 cuts is 40 jobs. This is a build-it-yourself recipe, and the docs make no promise of seamless joins.
Sources
Related posts
More in Developers
- Remove filler words from a video by API: cut at word timestamps
Descript's API lists Remove Filler Words as an Underlord edit. the Sume docs list no such op; here is how to cut um and uh yourself from words[] and a Timeline.
- OpenAI response_format json_schema on a Sume scheduled run
Sume accepts an OpenAI-shaped response_format as an alias for output_schema on schedule runs. Sending both returns 400, and the schema must be strict.
- Retool Workflow webhook needs X-Workflow-Api-Key: use a relay
Retool Workflows authenticate webhooks with an X-Workflow-Api-Key header or query parameter. Sume documents no custom delivery headers: relay after verifying.
- Rough cut from a script by API: one scene per Timeline slot
Descript Quick Design splits a script into moments with changing visuals. With Sume you map each scene to a Timeline video slot over a voiceover spine yourself.
Written by Sume