Cut dead air before the first word: STT start time, $0.01 split
Read words[0].start from a Sume STT job, then split the recording from that second for $0.01. For a 3-minute take the whole fix is $0.04. A copyable request.

To cut the silence before a speaker starts, transcribe the take with Sume STT, read words[0].start from the result, and send a timeline audio split job with one range that starts a little before that second. STT is $0.01 per audio minute and the split is $0.01 per job, so a 3-minute recording costs 3 x $0.01 + $0.01 = $0.04.
What the STT result gives you
Sume's STT route returns text, language fields when available, and words[], each { word, start, end } in seconds from the start of the audio. The first word's start is where speech begins, and the last word's end is where it finishes. Those two numbers are all a trim needs.
Leave a lead-in. Words are timed at their start, and a hard cut on that second can clip a breath or the first consonant. A margin of 0.1 to 0.2 seconds is a common choice. If the first word starts at 1.84, begin the range at 1.7.
| Step | Job | Basis | Cost |
|---|---|---|---|
| Find the first and last word | STT 1.0 | 3 audio minutes x $0.01 | $0.03 |
| Cut the file | Timeline audio split, 1 range | flat per job | $0.01 |
| Total | $0.04 |
The split request
The split needs operation: "split", a top-level url that is this workspace's media.sume.com audio, and ranges[] of 1 to 20. A range is { start, end? }; if you omit end, it runs to the end of the file. Keep the default wav output if the file will be joined or drive lip-sync, because mp3 adds priming padding at every edge.
curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: trim-take-07" \
-d '{
"operation": "split",
"url": "https://media.sume.com/artifacts/artf_demo/take07.wav",
"ranges": [{ "start": 1.7, "end": 171.3 }]
}'Checking the cut
Do not trust the numbers blindly. After the split, transcribe nothing again; instead listen to the first second of the new file and to its last second. If the first consonant is clipped, move the start back 0.1 second and re-run. A re-run is another $0.01, so two attempts still cost less than 3 cents on top of the $0.03 transcript.
Trim the tail the same way: the last word's end plus 0.3 seconds keeps the final syllable's decay. A hard cut at end can sound abrupt, particularly on a voice that trails off. For a render, a short fade-out on the soundtrack hides the rest.
A free alternative inside a render
Choose between a split and source_in by who else needs the trimmed file.
- If the file only goes into one Timeline render, use
audio.source_inon a single spine instead. It starts the spine at that offset, and the output length is stillaudio.duration_seconds. No separate job is needed. source_inis not allowed together withaudio.parts[]; use it for a single file.- A split gives you a reusable file with its own URL, which you can pass to captions or to an avatar job. Use a split when more than one step will use the trimmed audio.
- Always set
Idempotency-Key; the split is async by default and returns 202 unless you passmode: "sync".
Sources
Related posts
More in Developers
- Deno fetch: 25 s Wan 3.0 clip costs $3.125, 20-minute deadline
Deno script: submit a 25-second Wan 3.0 clip to Sume at 720p ($3.125), poll with backoff, stop at a 20-minute deadline. Run with deno run -A.
- Deno.serve webhook receiver for Sume jobs: Web Crypto HMAC, rotation
Verify Sume job webhooks in Deno with Web Crypto: refuse an empty secret, accept any rotation entry, 300 s tolerance, return 204. Retry window is 370 s.
- Idempotency-Key from a body hash plus a take number (Node, Sume)
Build an Idempotency-Key for Sume /v1/videos from a SHA-256 of the body plus a take counter, so a retry replays and a re-roll pays. Node code.
- Sume STT returns 400 for diarize: a two-speaker fix at $0.60 an hour
Sume STT 1.0 fixes diarize and tag_audio_events server-side and rejects them with 400. For two speakers, transcribe each track and merge by word times.
Written by Sume