Audit narration before posting a Short: flag 'here we see' lines
YouTube is reported to discount voice-over that only describes the screen. Transcribe a clip with Sume video inspect and flag describing lines first.

You can check a Short's narration for describe-the-screen lines before posting by transcribing it with Sume's video inspect and scanning the sentences for openers such as 'here we see' and 'as you can see'. That is a heuristic, not a test YouTube applies. Search Engine Journal reports that YouTube, announcing the October 1 change, said remixes need commentary, analysis and storytelling, not 'VO descriptions of what's happening on screen' (read 2026-10-03), and a transcript scan is a cheap way to catch your own habit.
What the scan cannot do is decide originality. YouTube published no threshold, no dashboard and no start date, per the same report, and its own monetization page judges whole videos by whether they add commentary, a storyline or a distinct focus (read 2026-10-03). Use the scan as a prompt to rewrite, not as a pass-fail gate.
What the transcript gives you
Video inspect with transcribe: true runs Sume STT 1.0 on the clip's audio. The public rate is $0.01 per audio minute. Adding segmentation: {"mode": "sentence"} returns gapless sentence segments[], and each has text, start and end. duration_seconds is a hint up to 600 and the source limit is 1,800 seconds. Check probe.has_audio first, because a silent clip returns inspect_source_has_no_audio.
The default call waits up to 30 seconds and answers with the finished inspect, or answers 202 with a job you poll. Off-host URLs are rejected, so the clip must be on media.sume.com first.
| Signal | How to compute it | What it suggests |
|---|---|---|
| Describing openers | Match 'here we see', 'as you can see', 'now we have' | The line reports the screen |
| Talk ratio | Sum of segment durations over clip length | Mostly silent, or mostly narrated |
| First-claim delay | Start time of the first sentence with 'because' or 'so' | How late the reasoning arrives |
| Repeated sentences | Same segment text more than once | A template or a loop |
A scan you can run
The function works on the segments list, so you can test it with a hand-made list before spending anything. Replace the sample with the transcript segments from your inspect result.
import re
PATTERNS = re.compile(
r"\b(here we see|as you can see|now we have|you can see)\b", re.I)
def audit(segments, clip_seconds):
flags = [s for s in segments if PATTERNS.search(s["text"])]
spoken = sum(s["end"] - s["start"] for s in segments)
return {"flagged": [(s["start"], s["text"]) for s in flags],
"talk_ratio": round(spoken / clip_seconds, 2)}
sample = [
{"text": "Here we see the settings page.", "start": 0.0, "end": 2.1},
{"text": "Turn this off because it slows the export.",
"start": 2.4, "end": 5.0},
]
print(audit(sample, 8.0))What to do with the flags
This is an audit for your own clips. It does not guarantee reach, does not predict what YouTube will recommend and says nothing about the accuracy of the claims in the narration.
- Rewrite each flagged line to say why, not what: 'this toggle slows the export' instead of 'here we see a toggle'.
- If most lines are flagged, the clip may need a different structure, not edits.
- If the talk ratio is very low, check that the captions carry reasoning.
- Rerun on the revised cut; a re-transcription is billed per audio minute.
Limits of a regex
A regex finds phrases, not meaning. 'Here we see a toggle that slows the export' contains the flagged opener but also a reason; the scan will still flag it. The reverse happens too: a line that describes the screen in different words will pass. Treat the output as a list of lines to read again, not a score.
Different languages need different patterns. Sume's STT takes a language_code hint such as en or ko, but the pattern list here is English, and I did not test the scan on other languages.
The scan also cannot judge the visuals. A Short with strong narration over a thin recording can still read as low effort, and a Short with sparse narration can be excellent if the visuals carry the lesson.
Cost and workflow
A 45-second clip is under a minute of audio, so the transcript line is about a cent at the documented public rate, plus the inspect job's compute, which the doc says is captured at actual container seconds. Run the scan on the cut you intend to publish, once, after the last edit. If you rewrite, re-record the voice track, then re-run on the final.
A second pass: repeated lines across a batch
Run the same scan across a whole batch and compare the segments. If five Shorts share the same closing sentence or the same second line, you have found the template a viewer would also notice. Collect all segment texts, count duplicates, and rewrite the ones that repeat. That one comparison is more useful than any single clip's score.
Sources
Related posts
More in Developers
- How many video jobs can I send now? The in-flight budget, worked
Concurrency 100, queue 500, 30 processing and 10 queued gives a new in-flight budget of 60 on Sume. Why the 450 wave hint is a hint, not your limit.
- No queue position or ETA for a Sume video job: what to show users
Sume exposes queue counts and remaining capacity, not a per-job queue position or ETA, and has no queue expiry option. What a video UI can say honestly instead.
- Node 22.23.3: pace Sume status polls with ratelimit headers
Read ratelimit-remaining, ratelimit-reset and retry-after from every Sume response in Node 22.23.3 and slow a wave of job polls before a 429 appears.
- Node 22.23.3 fetch: retry a Sume submit with one Idempotency-Key
Node 22.23.3 LTS bundles Undici 6.28.1. A fetch submit loop that retries 429 and 5xx with one Idempotency-Key so a retry never bills twice.
Written by Sume