TikTok Shop allows AI dubbing: localize a product video

TikTok Shop's AI policy lists AI-assisted translation and dubbing as permitted. Disclose synthetic voices, and note LIVEs ban AI voices. Batch per language.

4 min readSume
All posts

Yes. TikTok Shop's AI-generated content policy lists "AI-assisted translation and dubbing" and copywriting support among the permitted uses. Two cautions travel with it: a synthetic voice is a disclosure trigger, and the separate LIVE requirements prohibit AI-generated voices in LIVEs altogether.

One scope note about Sume. The docs pages used for this post do not describe a dubbing endpoint, so this post covers the parts Sume documents: per-language captions and batch fan-out.

What each TikTok page says

TikTok Shop Seller University pages, read 2026-10-05
QuestionAnswerPage
Is AI translation and dubbing allowed?Listed as permittedAI-generated content policy
Is copywriting support allowed?Listed as permittedAI-generated content policy
Do synthetic voices need disclosure?Yes: synthetic faces, voices or digital humans are a triggerAI-generated content policy
Are AI voices allowed in a LIVE?No: AI-generated voices and pre-recorded audio are prohibited in LIVEsHigh-quality videos and LIVEs
Must product info stay accurate?Yes: AI that creates false impressions about products is prohibitedAI-generated content policy

A pipeline for N languages

Keep one master edit and treat each language as a child job. Per language, the steps are: write or translate the script, produce the voice track, then burn captions in that language.

For the caption step, video captions accepts script_text, which keeps the speech-to-text word timings as the timing source and aligns your text to them. Its language field is only a speech-to-text hint (ko, en and so on) and never selects the style or font. If you author the text yourself, send cues with text, start and end and Sume burns it without speech recognition.

Each standalone caption job is $0.20 for clips up to 60 seconds, so 5 languages of one 45-second clip is 5 x $0.20 = $1.00 in caption jobs. That figure excludes the voice and the master clip.

Fan out with bulk runs

If the master edit is a Format, bulk runs queue up to 100 ordinary Format runs with a concurrency window of 1 to 16. Each item is the same body as a single run, so each language can carry its own instruction and its own spend cap. The queue has no webhook; poll GET /v1/format-run-queues/{id}, and branch on counts.failed because completed only means every item is terminal.

Quality gates per language

Translation can change a claim. A sentence that is accurate in one language can promise more in another, and the policy prohibits AI that misleads or creates false impressions about products. Have a speaker of each language read the script before the voice and the captions are produced, and check numbers, units and prices in the burned-in text.

Check timing as well. Caption jobs with script_text align your text to the speech-to-text timings and can fail with script_alignment_mismatch or script_alignment_failed, in which case the docs recommend simplifying the script or omitting script_text. Plan for one retry path per language.

Disclosure per language

A dubbed voice that is synthetic needs the same disclosure in every language version. Put the toggle or on-screen note on your posting checklist for each file, not once per product.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume