YouTube auto dubbing: 120-minute limit and no edits, so what now

YouTube dubs videos up to 120 minutes and will not let you edit a dub. When a clip is ineligible or wrong, here is the Sume pipeline for dubbing it yourself.

5 min readSume
All posts

YouTube's automatic dubbing is on by default for eligible creators, but a video over 120 minutes, one with little speech, an unsupported source language, fast speech or a copyright claim is not eligible, and you cannot edit a dub. When you need a dub YouTube will not make or one you can control, run the steps yourself: transcribe with video inspect, translate, synthesize speech, and render a new MP4 with Sume's Timeline.

The YouTube facts come from its Use automatic dubbing help page, read on 2026-10-02. Sume's side comes from video inspect, audio detach, Timeline 1.0 and MCP tools and gates. Sume has no one-call dubbing endpoint.

When does YouTube not dub a video?

The help page lists the reasons a video becomes ineligible. Each one is a case where doing it yourself is the only route.

The page also says creators cannot directly edit automatic dubs. If a translation is wrong, you can correct the original video language and the dubs regenerate. You can preview, publish, unpublish or delete generated dubs in YouTube Studio on desktop.

The page does not give a fix for a wrong source language beyond changing the original video language, so check the language setting of every upload before you trust a batch of dubs. A dub is only as good as the transcript under it, and rapid speech is the case where that transcript is most likely to slip.

YouTube automatic dubbing, ineligible and fixed points (read 2026-10-02)
SituationWhat YouTube saysDub it yourself?
Video over 120 minutesIneligibleYes, in pieces
Minimal speechIneligibleLittle to dub
Unsupported source languageIneligibleYes, if you can translate it
Rapid speech pacingIneligibleYes, with a new pacing
Copyright claimIneligibleCheck rights first
Wrong translationNo direct edit; fix the original language and regenerateYes, you control the script

What does a do-it-yourself dub look like on Sume?

There are four stages, and Sume covers three of them as separate steps. Detach or inspect gives you the speech and its timing. Translation is your job or an Agent Completion's. Text-to-speech is the tts_create tool in Sume's hosted MCP, and Timeline joins the result onto the picture.

Start with POST /v1/video-inspect and transcribe: true, with language_code when you know the source language. The result carries the transcript text, word timings and sentence segments. Probe and stills are unbilled; transcription has its own cost, so read the catalog.

Translation is the step where you gain control. Because you hold the text, you can keep product names, brand terms and numbers fixed, which an automatic dub does not let you edit. Review the translated script against the source line by line before you synthesize anything; a mistake caught as text costs nothing, and one caught as audio costs a render.

curl -X POST https://api.sume.com/v1/video-inspect \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: dub-es-inspect-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
    "frames": false,
    "transcribe": true,
    "language_code": "en"
  }'

What are the real limits?

Sume's media tools read a source of at most 1800 seconds and produce at most 900 seconds of trim or detach output; Timeline renders up to 1800 seconds. A 120-minute video is 7200 seconds, so you split it first, dub each part and join the parts. That is covered in the podcast split plan.

There is no lip-sync step. Sume's models page says video models do not lip-sync to generated speech, so a dubbed talking head will not match mouth movement. Cutaway-heavy, voiceover-led or screen-recording videos dub well this way; a close-up interview does not.

Speech takes different time in different languages. A Spanish line is often longer than the English line it replaces, so the new audio may run past its picture. Timeline sets output length from the audio spine and lets a video slot trail it by at most 0.5 seconds, so plan first with the unbilled plan call, then shorten the script or add a held shot where the speech runs long.

How does a self-made dub reach YouTube?

The self-made dub is a normal MP4, not a YouTube audio track. Upload it as a separate video in the target language, or as a new language version of your channel. Sume does not upload to YouTube or add an alternate audio track to an existing upload, so check what your Studio account offers for adding audio.

Disclosure still matters: see turning off YouTube auto dubbing and disclosure before you publish synthetic speech.

Whether you dub yourself or let YouTube do it, listen to the first and last minute of each dub in the target language, ideally with a native speaker, before you publish.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume