YouTube Studio draft feedback: measure pacing yourself with an API
YouTube is adding Studio draft feedback on pacing and structure. Sume cannot copy it, but video inspect gives you word rate and stills to measure pacing.

The Made on YouTube announcements say Studio will give feedback on drafts covering pacing, structure, storytelling and more, on web and mobile. That feedback lives inside YouTube; Sume does not read it or reproduce it. What you can do before upload is measure the parts of pacing that are numbers: speech rate, silence and how often the picture changes.
What is stated, and what you can measure
The YouTube posts name the feedback categories but do not describe a scoring method or thresholds, so there is no target to hit. Treat your own numbers as comparisons between drafts of the same video.
| Topic YouTube names | What you can measure | Sume call |
|---|---|---|
| Pacing | Words per minute, pauses | video inspect with transcribe true |
| Structure | Where the hook, middle and ending fall | Stills at fixed times, frames.at |
| Storytelling | Not measurable by a probe | Review the stills yourself |
Pull the transcript and stills in one call
Video inspect takes a Sume-hosted clip of up to 1800 seconds. With transcribe: true it adds Sume STT at $0.01 per audio minute, and segmentation.mode: sentence returns sentence segments. frames: {fps: 0.5} returns a still every 2 seconds, up to 24 per call.
curl -X POST https://api.sume.com/v1/video-inspect \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: draft-pacing-001" \
-d '{"video_url":"https://media.sume.com/artifacts/artf_demo/draft1.mp4","transcribe":true,"segmentation":{"mode":"sentence"},"frames":{"fps":0.5}}'Turn it into a comparison
Words per minute is the transcript word count divided by the clip minutes. Run the same call on draft 1 and draft 2, then compare. The stills show whether the first 10 seconds reach the point; they also show whether the visual changes at least every few seconds, which is a judgment you make from the sheet rather than something the API scores.
If a draft is too slow, cut dead air with video trim and measure again. Keep the inputs identical so the comparison is fair.
Sources
Related posts
More in Use cases
- YouTube thumbnail 1920x1080 is not a legal GPT Image 2.5 size
1080 is not a multiple of 16, so Sume's custom size rule rejects 1920x1080 for GPT Image 2.5. Use 1280x720, 2560x1440 or 3840x2160 for 16:9 thumbnails.
- 36 thumbnail test images: list-price totals across 7 image models
A 12-concept, 3-variant thumbnail test is 36 images. Fal list prices of $0.03 to $0.15 give $1.08 to $5.40; Sume's endpoint pricing line is what you pay.
- YouTuber's year-in-review video: 10 clips, 6 thumbnails, $1.18 on Sume
A 3-minute year-in-review cut from ten past videos costs $1.175 on Sume, including ten trims, audio detaches, a render, a music bed and six 16:9 thumbnails.
- Zillow Preview teaser video: a coming-soon reel from photos
Zillow Preview listings are now live on Realtor.com too. Build an 18-second vertical coming-soon teaser from listing photos with Sume Timeline and caption cues.
Written by Sume