audio_source_in_requires_single_spine: source_in needs one voice file
audio.source_in only works with one audio.url. With audio.parts, set source_in on each part instead; the Sume hint is use_part_source_in.

audio_source_in_requires_single_spine means you sent audio.source_in next to audio.parts[]. A spine-level in-point would not say which slice it seeks into, so Sume refuses it. Either use a single audio.url with audio.source_in, or keep parts[] and give each part its own source_in. The next_action is use_part_source_in.
The use case is trimming dead air at the start of a take. A recording that opens with three seconds of breath is a poor hook for a vertical clip, and YouTube lists Shorts as vertical video up to 3 minutes (YouTube Help, read 2026-10-05), so every second counts.
Single spine form
With one url, source_in is the in-point into that file. The output length is still duration_seconds, so the spine reads from source_in for that many seconds. This is described in the Timeline 1.0 docs program table.
{
"audio": {
"url": "https://media.sume.com/artifacts/artf_demo/take.wav",
"source_in": 3.2,
"duration_seconds": 40
},
"video": [
{ "source_url": "https://media.sume.com/artifacts/artf_demo/a.mp4", "start": 0, "duration": 40 }
]
}Parts form
With parts[], each slice is a url plus optional source_in and duration. Slices join gaplessly in the sample domain, with no re-generation of speech. If you give every part a duration, the sum must reach duration_seconds, otherwise audio_parts_shorter_than_duration is returned.
| Shape | audio.source_in | parts[].source_in |
|---|---|---|
| audio.url only | allowed | not applicable |
| audio.parts only | audio_source_in_requires_single_spine | allowed |
| audio.mode silence | silent_audio_takes_no_source_in | not applicable |
| video[].source_in | separate field: in-point into the clip | separate field |
Do not confuse the two source_in fields
video[].source_in is the in-point into a video file and has nothing to do with the audio rule. A common mistake is moving a video in-point into the audio object when the voice and the footage came from the same camera file. If the footage and voice are one file, audio detach can make the voice a separate wav, then the two in-points are set independently.
Picking the in-point
Find the in-point by looking at the waveform of the take, or by running a transcript and using the first word time. Sume's video inspect can return word times when transcribe: true is set, at the public rate of $0.01 per audio minute, and the first word start is a good source_in minus a short breath of about a tenth of a second. Treat that margin as a starting guess, not a rule. Cutting exactly on the first consonant can clip it, so listen to the result once before publishing.
Limits and a safe workflow
Run /plan to catch the code before spending anything; it validates the schema and runs the compiler without a job. Remember that all URLs must be your own media.sume.com artifacts, so import the take first. If you only need to cut a clip to a range and keep it as a file, video-trim makes a new MP4 for $0.02 per job, but for a voice-only edit the in-point approach above avoids an extra file.
Sources
Related posts
More in Media tools
- Avatar clip frame rate: Griffin 25 fps vs Fabric vs H3 Max
Tavus says Griffin streams 8-frame latents at 25 fps. Sume's docs say Fabric is 25 fps and H3 Max must be measured. Probe a clip with video-inspect.
- Baby shower video music: a gentle 90-second slideshow bed
Pick a gentle instrumental for a 90-second baby shower photo slideshow: one generation and a two-minute render, $0.325 on Sume.
- Put a 4:5 AI image on a 9:16 canvas with blurred fill in Pillow
Fill the empty bands of a 9:16 canvas with a blurred, darkened copy of the same Sume image, and place the sharp 4:5 original over it. Code and blur settings.
- Burn captions: pick one of script_text, words, cues or segments
Sume video captions accepts one wording source per job: script_text, words, cues or segments. When each fits and which ones skip speech-to-text.
Written by Sume