Audit narration before posting a Short: flag 'here we see' lines

YouTube is reported to discount voice-over that only describes the screen. Transcribe a clip with Sume video inspect and flag describing lines first.

5 min readSume
All posts

You can check a Short's narration for describe-the-screen lines before posting by transcribing it with Sume's video inspect and scanning the sentences for openers such as 'here we see' and 'as you can see'. That is a heuristic, not a test YouTube applies. Search Engine Journal reports that YouTube, announcing the October 1 change, said remixes need commentary, analysis and storytelling, not 'VO descriptions of what's happening on screen' (read 2026-10-03), and a transcript scan is a cheap way to catch your own habit.

What the scan cannot do is decide originality. YouTube published no threshold, no dashboard and no start date, per the same report, and its own monetization page judges whole videos by whether they add commentary, a storyline or a distinct focus (read 2026-10-03). Use the scan as a prompt to rewrite, not as a pass-fail gate.

What the transcript gives you

Video inspect with transcribe: true runs Sume STT 1.0 on the clip's audio. The public rate is $0.01 per audio minute. Adding segmentation: {"mode": "sentence"} returns gapless sentence segments[], and each has text, start and end. duration_seconds is a hint up to 600 and the source limit is 1,800 seconds. Check probe.has_audio first, because a silent clip returns inspect_source_has_no_audio.

The default call waits up to 30 seconds and answers with the finished inspect, or answers 202 with a job you poll. Off-host URLs are rejected, so the clip must be on media.sume.com first.

Narration signals a script can compute, read 2026-10-03
SignalHow to compute itWhat it suggests
Describing openersMatch 'here we see', 'as you can see', 'now we have'The line reports the screen
Talk ratioSum of segment durations over clip lengthMostly silent, or mostly narrated
First-claim delayStart time of the first sentence with 'because' or 'so'How late the reasoning arrives
Repeated sentencesSame segment text more than onceA template or a loop

A scan you can run

The function works on the segments list, so you can test it with a hand-made list before spending anything. Replace the sample with the transcript segments from your inspect result.

import re

PATTERNS = re.compile(
    r"\b(here we see|as you can see|now we have|you can see)\b", re.I)

def audit(segments, clip_seconds):
    flags = [s for s in segments if PATTERNS.search(s["text"])]
    spoken = sum(s["end"] - s["start"] for s in segments)
    return {"flagged": [(s["start"], s["text"]) for s in flags],
            "talk_ratio": round(spoken / clip_seconds, 2)}

sample = [
    {"text": "Here we see the settings page.", "start": 0.0, "end": 2.1},
    {"text": "Turn this off because it slows the export.",
     "start": 2.4, "end": 5.0},
]
print(audit(sample, 8.0))

What to do with the flags

This is an audit for your own clips. It does not guarantee reach, does not predict what YouTube will recommend and says nothing about the accuracy of the claims in the narration.

  • Rewrite each flagged line to say why, not what: 'this toggle slows the export' instead of 'here we see a toggle'.
  • If most lines are flagged, the clip may need a different structure, not edits.
  • If the talk ratio is very low, check that the captions carry reasoning.
  • Rerun on the revised cut; a re-transcription is billed per audio minute.

Limits of a regex

A regex finds phrases, not meaning. 'Here we see a toggle that slows the export' contains the flagged opener but also a reason; the scan will still flag it. The reverse happens too: a line that describes the screen in different words will pass. Treat the output as a list of lines to read again, not a score.

Different languages need different patterns. Sume's STT takes a language_code hint such as en or ko, but the pattern list here is English, and I did not test the scan on other languages.

The scan also cannot judge the visuals. A Short with strong narration over a thin recording can still read as low effort, and a Short with sparse narration can be excellent if the visuals carry the lesson.

Cost and workflow

A 45-second clip is under a minute of audio, so the transcript line is about a cent at the documented public rate, plus the inspect job's compute, which the doc says is captured at actual container seconds. Run the scan on the cut you intend to publish, once, after the last edit. If you rewrite, re-record the voice track, then re-run on the final.

A second pass: repeated lines across a batch

Run the same scan across a whole batch and compare the segments. If five Shorts share the same closing sentence or the same second line, you have found the template a viewer would also notice. Collect all segment texts, count duplicates, and rewrite the ones that repeat. That one comparison is more useful than any single clip's score.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume