TikTok reservation ads must include audio: probe, then add a bed
TikTok's reservation spec says every video creative must include audio. Probe has_audio with video-inspect, then give a silent clip a voice or music spine.

A silent clip is a failed reservation creative. TikTok's reservation in-feed page says all video creatives must include audio (read 2026-10-07). Before upload, probe the file with Sume's video-inspect and read has_audio; if it is false, assemble the clip again with a voice or music spine in Timeline 1.0.
The check is cheap and the fix is one render. The part to avoid is declaring silence: a Timeline audio.mode: "silence" produces a declared-length track with no file, which is the opposite of what the page asks for.
What the page says
The same reservation page lists 5-60 seconds with 9-15 recommended, a 500 MB cap, and a 2,500 kbps floor. It also says watermarks, including TikTok's own branding, are prohibited. The audio line is short; it does not name a loudness target, a codec, or a minimum track length, so I do not invent one.
| Item | What the page says |
|---|---|
| Audio | all video creatives must include audio |
| Duration | 5-60 s, recommend 9-15 s |
| Watermarks | prohibited, including TikTok's own branding |
| Preview | use TikTok's preview tool to check devices |
Step 1: probe with frames off
Send the clip URL with frames: false. The clip must already be a media.sume.com artifact or asset of your workspace; import it first with POST /v1/media-imports, per Media inputs. Read probe.has_audio in the answer.
curl -X POST https://api.sume.com/v1/video-inspect \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: audio-probe-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/clip.mp4",
"frames": false
}'Step 2: give the clip a spine
Timeline 1.0 takes one audio spine plus ordered video[] slots. For a 12-second ad, send audio.url (a voiceover or a track you made, hosted on Sume) with audio.duration_seconds: 12, and place the clip in video[0] with start: 0. Optionally add a soundtrack bed with duck_db so the music dips under speech; ducking needs a real spine, not silence.
Music 1.0 makes a track from a text prompt. It does not take a duration field, so ask for the length in the prompt and trim in the timeline. The public Timeline rate is $0.10 per ceil output minute, so a 12-second ad reserves one minute (confirm in GET /v1/catalog).
Limits
Sume does not inspect your account or submit anything to TikTok. A passing has_audio only proves a track exists; whether the sound is acceptable to reviewers is TikTok's call. Also watch the audio you inherit from a downloaded clip: the watermark rule above applies to the pictures you reuse.
A quick decision list
If has_audio is true and the track is the voice you want, upload. If it is true but the audio came from a stock clip you cannot clear, replace it. If it is false, build the spine first, then the picture.
Mind sync when you add a spine to existing footage: the timeline places slots on the spine's clock, and video[0].start must be 0, so cut the picture to the voice rather than the other way round. A 12-second voice track means 12 seconds of slots, with coverage allowed to stop at most 0.5 seconds early.
Sources
Related posts
More in Media tools
- TopView says 15 seconds, reservation says 9-15: trim both for $0.04
TikTok's TopView page recommends 15 seconds and the reservation page 9-15. Cut a 9-second and a 15-second version from one master with two video-trim jobs.
- TikTok TopView locks assets at preloading: freeze the cut first
TopView creatives need sales-rep pre-approval and cannot change once preloading begins. Use Timeline's unbilled plan, then render once with an idempotency key.
- Do fades shift my cuts? Timeline start times with transitions
In Timeline 1.0 a slot's declared start is its place on the audio spine, fade or no fade. The compiler compensates for the xfade overlap; gaps hold a frame.
- Two-voice dialogue audio on Sume: TTS lines joined with audio concat
Build a role-play or interview track by generating one TTS line per turn in each speaker's voice, then joining up to 20 turns into one gapless file for $0.01.
Written by Sume