TikTok AI Outline splits a video in six: plan six slots in Sume first

TikTok's AI Outline breaks a video into six parts. Mirror that outline in Sume with six timeline slots and the unbilled plan call, then render once.

5 min readSume
All posts

Short answer

TikTok's AI Outline is an in-app tool that breaks a video into six parts, and Sume does not read or call it. If you want the same six-part shape for a clip you control, write the outline yourself as six Timeline 1.0 slots and check it with the unbilled plan call before you pay for a render.

That gives you the part boundaries in a file you can review, version and rerun, instead of a layout that lives only inside an app.

What TikTok says about AI Outline

TikTok's newsroom post on its new AI-powered tools describes AI Outline as a tool that breaks a video into six parts (read 2026-10-05). The same page says it launched on Oct 28, 2025, that it is for users 18 and older, and that it is available in the US, Canada and select markets. I did not find a statement about an API for it, so treat it as an app feature.

The post also describes Smart Split, which works on a source video over 1 minute and clips, reframes and captions it in TikTok Studio on the web. Smart Split and AI Outline solve different problems, and this post is only about the six-part outline.

The facts side by side

Each row is a fact from a page I read, or from the Sume docs in the repo.

AI Outline and Timeline 1.0 plan, read 2026-10-05
ItemValueSource
AI Outline outputA video broken into six partsTikTok Newsroom
AI Outline launchOct 28, 2025TikTok Newsroom
AI Outline audience18+, US, Canada and select marketsTikTok Newsroom
Timeline slots per render1 to 200 in video[]Sume docs
Plan callPOST /v1/timeline-1.0/plan, unbilled, no job createdSume docs
Render rate$0.10 per ceil(output minute)Sume docs
Plan fields returnedduration_seconds, segment_count, billable_minutes, estimated_cost_usd_microsSume docs

Six slots in a plan request

This sketch cuts six 15-second parts from one source clip into a 90-second output. Each slot starts on the output timeline at a multiple of 15 and reads from its own in-point in the source through source_in. The first slot must start at 0. The audio spine uses mode silence with a declared length, so no voice file is needed for the plan. Replace the URL with a clip you imported to media.sume.com first, because the API refuses off-host URLs.

import json, os, urllib.request

key = os.environ.get("SUME_API_KEY", "")
if not key:
    raise SystemExit("set SUME_API_KEY")
src = "https://media.sume.com/artifacts/artf_demo/talk.mp4"
slots = [
    {"source_url": src, "start": i * 15, "duration": 15, "source_in": i * 40}
    for i in range(6)
]
body = {"audio": {"mode": "silence", "duration_seconds": 90}, "video": slots}
req = urllib.request.Request(
    "https://api.sume.com/v1/timeline-1.0/plan",
    data=json.dumps(body).encode(),
    headers={"Authorization": "Bearer " + key,
             "Content-Type": "application/json"},
)
print(json.load(urllib.request.urlopen(req)))

Read the plan before you render

Six slots of 15 seconds make 90 seconds, which is 2 billable minutes because the rate rounds the output up to whole minutes. At $0.10 per minute that is $0.20, so estimated_cost_usd_micros should read 200000 if the live rate matches the docs. Confirm the live rate in GET /v1/catalog. The plan does not download media, so it cannot predict padding or looping warnings from sources shorter than their slots.

If the numbers match, send the same body to POST /v1/timeline-1.0/render with an Idempotency-Key header, then poll the job. Six segments is under the 12-segment point where the strategy starts to chunk, so there is nothing to tune.

Limits to keep in mind

  • Hard cuts are the default. A fade between parts is optional, and more than 8 chained fades is refused, so insert a cut after a run of eight.
  • Stills are static holds, so an outline built from images will not pan or zoom.
  • The outline is yours. Sume does not detect topic changes in the source, so pick the in-points from the transcript or from stills (video inspect gives both).

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume