Flag near-identical openings in your last 20 Short scripts
A short Python script that compares the first sentence of each Short script and flags pairs that read alike, a cheap check against templated AI Shorts.

YouTube's monetization page, read on October 5, 2026, describes AI-generated content made with generic or unoriginal templates, giving the impression of mass production. It gives no similarity threshold, and 9to5Google reports that the October 2 recommendation update draws no clear line either. You can still measure your own sameness. The easiest signal is the opening line.
The idea
Keep each Short's script as a text file. Take the first sentence of each, compare every pair with difflib's SequenceMatcher, and print the pairs above a ratio. The threshold is yours; 0.6 is a reasonable start for short sentences because swapping one noun in a seven-word line already gives about 0.85.
This is a heuristic. It cannot tell you whether the idea is original. It catches the failure you can fix in minutes: the same sentence with a new topic.
The script
Put the scripts in a folder called scripts, one .txt per Short, and run it with Python 3.
import glob, re
from difflib import SequenceMatcher
LIMIT = 0.6
def opening(text):
text = " ".join(text.split())
return re.split(r"(?<=[.!?])\s", text, maxsplit=1)[0].lower()
openings = {}
for path in sorted(glob.glob("scripts/*.txt"))[-20:]:
with open(path, encoding="utf-8") as f:
openings[path] = opening(f.read())
names = list(openings)
flagged = 0
for i in range(len(names)):
for j in range(i + 1, len(names)):
a, b = openings[names[i]], openings[names[j]]
score = SequenceMatcher(None, a, b).ratio()
if score >= LIMIT:
flagged += 1
print(round(score, 2), names[i], "|", names[j])
print(" ", a)
print(" ", b)
print(flagged, "pairs at or above", LIMIT)What to do with a flagged pair
Rewrite the later opening, not the earlier one, so a published Short is not edited after the fact. Change the structure of the line, not just a word: swap a statement for a question, a number for a scene, a claim for a quoted objection.
Then extend the idea. The same loop works on the last sentence, which is where templated calls to action repeat. Add a second check on word count per script, since near-equal lengths are another sign of one frame reused.
Limits of the check
Two Shorts can open differently and still tell the same story. Two can open alike on purpose, as a series signature, which the monetization page tolerates for intros and outros. Use the output as a prompt for a human decision.
Run it on a schedule
Add the check to the step that writes the script, not the step that publishes. If a new opening scores above the limit against any of the last 20, the writer, human or model, gets the flagged pair back and tries again. That makes the check part of drafting, which is cheaper than a failed review later.
Keep the threshold in one place and log the scores. After a month you will see whether your channel's range of openings is widening or collapsing, which is more useful than any single score.
Sources
Related posts
More in Developers
- Python and SQLite: log every video job's usage.cost by model
A 13-line stdlib script stores each completed job's usage.cost in SQLite, ignores duplicates by job id and prints spend per model. Seedance 2.5 5 s is $2.89.
- Python receiver for a Sume /v1/videos callback_url, signature checked
A standard-library Python webhook receiver for Sume video jobs: checks x-sume-webhook-signature, refuses an empty secret, rejects stale timestamps.
- Python Sume webhook handler that accepts the webhook.test event
Verify the sume-v1 signature over timestamp.body, refuse an empty secret, and accept webhook.test, which has no job_id. Stdlib Python, runs offline.
- Python: three Wan 3.0 hooks from one reference image, with costs
A Python script for Sume's /v1/videos: submit three Wan 3.0 hook prompts with one reference image at 480p, poll each job, and print the usage cost.
Written by Sume