Cut dead air before the first word: STT start time, $0.01 split

Read words[0].start from a Sume STT job, then split the recording from that second for $0.01. For a 3-minute take the whole fix is $0.04. A copyable request.

5 min readSume
All posts

To cut the silence before a speaker starts, transcribe the take with Sume STT, read words[0].start from the result, and send a timeline audio split job with one range that starts a little before that second. STT is $0.01 per audio minute and the split is $0.01 per job, so a 3-minute recording costs 3 x $0.01 + $0.01 = $0.04.

What the STT result gives you

Sume's STT route returns text, language fields when available, and words[], each { word, start, end } in seconds from the start of the audio. The first word's start is where speech begins, and the last word's end is where it finishes. Those two numbers are all a trim needs.

Leave a lead-in. Words are timed at their start, and a hard cut on that second can clip a breath or the first consonant. A margin of 0.1 to 0.2 seconds is a common choice. If the first word starts at 1.84, begin the range at 1.7.

Dead-air trim for a 3-minute recording, rates read 2026-10-09
StepJobBasisCost
Find the first and last wordSTT 1.03 audio minutes x $0.01$0.03
Cut the fileTimeline audio split, 1 rangeflat per job$0.01
Total$0.04

The split request

The split needs operation: "split", a top-level url that is this workspace's media.sume.com audio, and ranges[] of 1 to 20. A range is { start, end? }; if you omit end, it runs to the end of the file. Keep the default wav output if the file will be joined or drive lip-sync, because mp3 adds priming padding at every edge.

curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: trim-take-07" \
  -d '{
    "operation": "split",
    "url": "https://media.sume.com/artifacts/artf_demo/take07.wav",
    "ranges": [{ "start": 1.7, "end": 171.3 }]
  }'

Checking the cut

Do not trust the numbers blindly. After the split, transcribe nothing again; instead listen to the first second of the new file and to its last second. If the first consonant is clipped, move the start back 0.1 second and re-run. A re-run is another $0.01, so two attempts still cost less than 3 cents on top of the $0.03 transcript.

Trim the tail the same way: the last word's end plus 0.3 seconds keeps the final syllable's decay. A hard cut at end can sound abrupt, particularly on a voice that trails off. For a render, a short fade-out on the soundtrack hides the rest.

A free alternative inside a render

Choose between a split and source_in by who else needs the trimmed file.

  • If the file only goes into one Timeline render, use audio.source_in on a single spine instead. It starts the spine at that offset, and the output length is still audio.duration_seconds. No separate job is needed.
  • source_in is not allowed together with audio.parts[]; use it for a single file.
  • A split gives you a reusable file with its own URL, which you can pass to captions or to an avatar job. Use a split when more than one step will use the trimmed audio.
  • Always set Idempotency-Key; the split is async by default and returns 202 unless you pass mode: "sync".

Sources

Related posts

More in Developers

All Developers posts

Written by Sume