YouTube description limit is 5000 bytes, not characters: check it

The Data API caps snippet.description at 5000 bytes and snippet.title at 100 characters. A Python check for both, and where a Sume transcript fits.

5 min readSume
All posts

The Data API caps snippet.description at 5000 bytes, not 5000 characters, and snippet.title at 100 characters. A description of plain English text hits the limit at about 5000 characters, but one with accented letters, emoji or non-Latin script hits it sooner, because those characters take more than one byte in UTF-8.

The limits come from the YouTube Data API Videos resource, read on 2026-10-03: the title "may contain all valid UTF-8 characters except < and >" with a maximum of 100 characters, and the description has the same character rule with a maximum of 5000 bytes. Sume does not write to YouTube, so the check below runs on your side before you call videos.insert.

Why do bytes matter if my text is English?

Two reasons. First, a transcript pasted into a description often carries curly quotes, long dashes and arrows, each of which is three bytes in UTF-8 even though it looks like one character. Second, a Short with a dubbed version needs a description in another script, where nearly every character is multi-byte.

The page's wording is plain: bytes for the description, characters for the title. A script that counts len(text) for both will pass a description that the API then rejects.

Limits on the Videos resource (read 2026-10-03)
FieldLimitUnit
snippet.title100characters
snippet.description5000bytes
snippet.tags[]500characters for the whole list

How do I check both in Python?

This script counts the title in characters and the description in UTF-8 bytes, rejects the two forbidden angle-bracket characters, and trims an over-long description at a character boundary so it never cuts a multi-byte character in half. It runs as-is.

def check(title, description):
    problems = []
    if len(title) > 100:
        problems.append("title is %d characters (max 100)" % len(title))
    size = len(description.encode("utf-8"))
    if size > 5000:
        problems.append("description is %d bytes (max 5000)" % size)
    for text in (title, description):
        if "<" in text or ">" in text:
            problems.append("angle brackets are not allowed")
    return problems

def trim(description, limit=5000):
    out = ""
    for ch in description:
        if len((out + ch).encode("utf-8")) > limit:
            break
        out += ch
    return out

long_text = "caf\u00e9 " * 1300
print(check("Episode 3", long_text))
print(len(trim(long_text).encode("utf-8")))

Where can a Sume transcript help?

Video inspect can transcribe one media.sume.com clip with transcribe: true at a public rate of $0.01 per audio minute, and segmentation.mode: "sentence" also returns gapless sentence segments. That gives you text to build a description or chapter list from. Probe and stills are billed by their Modal compute, so confirm live pricing in GET /v1/catalog.

Use the transcript as raw material, not as the description. Trim it with the function above, keep the first lines for the viewer, and put links where they read well. Sume does not decide what YouTube shows from a description; the pages I read do not say how the field is used for ranking, so this post makes no claim about it.

Where do chapters and titles go?

Chapter lines live inside the description, so they spend the same byte budget; chapters from a transcript shows the layout. For the title side, the 100-character Shorts title check goes deeper, and sentence segmentation for cues shows how the sentence segments are shaped.

What should I do when the description is too long?

Cut from the bottom, not the top. The first lines are the ones a viewer sees before expanding the description, so put the answer and the most useful link there and let the trim function drop the tail. If you generated the description from a transcript, drop whole sentences rather than cutting mid-word; the sentence segments Sume returns make that straightforward.

Remember that the same limit applies to every language version. A description in Korean or Japanese uses three bytes for most characters in UTF-8, so a text of roughly 1,600 characters can already be at the cap. The title is the opposite: it counts characters, so a 100-character title in any script passes the title rule even though its bytes are many more.

Finally, run the check on the final string, after you have appended links, hashtags and credits. A check that runs on the draft and not on the assembled text is the usual way a batch fails on row 40.

  • Count the description in UTF-8 bytes and the title in characters.
  • Reject the angle-bracket characters in both.
  • Trim at a character boundary, never inside a byte sequence.
  • Run the check on the assembled text, not the draft.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume