TikTok video_cover_timestamp_ms: pick the cover frame from a clip
TikTok takes a cover frame as a millisecond offset. Pull candidate stills with Sume video frames, choose one, and send its time as video_cover_timestamp_ms.

TikTok does not take a cover image in the Direct Post request. It takes a time: video_cover_timestamp_ms, the millisecond offset of the frame to use. If the value is invalid, the cover falls back to the first frame, which for a generated clip can be a soft or half-formed frame. To choose well, pull a handful of stills at known times with Sume's unbilled video frames endpoint, pick the best one, and send that time.
What TikTok documents
| Field | Meaning on the page |
|---|---|
video_cover_timestamp_ms | int32; the frame, in milliseconds, used as the cover |
| Invalid value | Defaults to the first frame |
title | Up to 2200 UTF-16 runes |
privacy_level | Must match an option from the creator info query |
Step 1: sample candidates
POST /v1/video-frames takes video_url (a clip in your workspace) and exactly one of at[] or fps. It returns 202 and a job id; read GET /v1/video-frames/:id until resource_status is ready. The clip must be on media.sume.com and at most 300 seconds, with 1 to 24 times per call. Frames come back at source size as durable image URLs, and the call is unbilled.
import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
clip = "https://media.sume.com/artifacts/artf_demo/talk.mp4"
times = [0.5, 1.5, 2.5, 3.5]
r = requests.post(API + "/v1/video-frames", timeout=60,
headers={**H, "Idempotency-Key": "cover-probe-001"},
json={"video_url": clip, "at": times})
r.raise_for_status()
fid = r.json()["data"]["request_id"]
while True:
d = requests.get(f"{API}/v1/video-frames/{fid}", headers=H,
timeout=60).json()["data"]
if d["resource_status"] in ("ready", "failed"):
break
time.sleep(2)
for f in d["frames"]:
print(int(f["t"] * 1000), f["url"])
Step 2: send the chosen time
Look at the printed URLs, pick one, and put its millisecond value in post_info.video_cover_timestamp_ms. A frame whose url is null failed to extract; that does not fail the job, so skip it. A practical rule is to avoid the first half second, where generated clips can start on a soft frame, and to prefer a time where the subject is in focus and any burned-in caption is fully visible. Four to eight candidates is plenty; the endpoint accepts a list, so one call covers them all.
Limits
The cover time must fall inside the file you actually upload. If you trim the clip afterwards with video trim, re-read the new actual_start_seconds and shift the time, or sample frames from the trimmed MP4 instead. The sample sets Idempotency-Key to a fixed string, so change it per clip. TikTok may re-encode the file, and the page does not promise the cover matches a Sume still pixel for pixel.
Sources
Related posts
More in Use cases
- Translated with Meta AI label: how viewers turn Reels dubbing off
Meta labels every translated Reel Translated with Meta AI. A viewer picks Don't translate in the audio and language section of the three-dot menu to opt out.
- Trim a long video then caption the clip: two Sume jobs in order
Cut a moment from a long recording with video-trim, then burn captions on the new MP4. Why the order matters, what each job costs, and which URL goes where.
- Upload 15 YouTube Shorts at a time: prepare the batch with Sume
YouTube's upload page lets you pick up to 15 Shorts at once. Prepare 15 correctly sized clips with Sume trim jobs, one idempotency key each.
- Walmart Sponsored Videos run 5 to 45 seconds: cut yours to length
Walmart Connect says Sponsored Videos run between five and 45 seconds. How to trim a finished product video to that window with Sume's video-trim endpoint.
Written by Sume