Use a Pocket TTS voiceover made outside Sume in a Sume video

Media imports reject audio. A local WAV needs the asset upload route, then a media.sume.com URL as Timeline audio.url. What I verified, and what I did not run.

5 min readSume
All posts

Sume's POST /v1/media-imports only takes public TikTok or Instagram video URLs, so it will not take a WAV from Kyutai Pocket TTS. The route that accepts audio is the asset upload flow: create an upload URL, PUT the bytes, complete the asset, then use the returned media.sume.com URL as audio.url in a Timeline render. That route is implemented but hidden from the public OpenAPI, so treat it as unstable.

I read Kyutai's Pocket TTS README and voices page, Sume's Timeline and Timeline audio docs, the Sume API reference and the API route code on 2026-10-03. I tested the script below against a mocked server, not live.

What do the docs say about getting audio in?

The Timeline docs say every URL must already be on media.sume.com and tell you to import first. The import endpoint, per the API reference, imports TikTok or Instagram videos at $0.15 each and rejects other hosts. The API notes list /v1/assets/upload-url and /v1/assets/:id/complete as implemented but hidden, and say not to treat them as public contract.

What do you need from Kyutai?

The README describes Pocket TTS as a 100M-parameter CPU model, about 6x real time on a MacBook Air M4, with seven languages. Its Python example writes a wav with the model's sample_rate. The repository LICENSE is an MIT-style permission notice, and the voices page lists a license per voice, so check the voice you used before publishing. If your file is 32-bit float, convert it to 16-bit PCM first; I have not tested float WAV on Sume. The Pocket TTS CLI also accepts a plain wav as the voice for cloning, which is a separate question from getting finished audio into a video.

Routes for outside audio on Sume, read 2026-10-03
RouteTakes audio?Public contract?
POST /v1/media-importsNo: TikTok and Instagram video onlyYes
POST /v1/assets/upload-url, PUT, /completeYes: media_type audioNo: hidden from OpenAPI
Timeline audio.urlOnly media.sume.com URLsYes

How do you upload and plan the video?

The script uploads voiceover.wav, completes it, then runs the unbilled Timeline plan with the returned URL. Set VOICEOVER_SECONDS to the clip length and CLIP_URL to a media.sume.com video. If the plan looks right, send the same body to /v1/timeline-1.0/render with an Idempotency-Key. Timeline lists $0.10 per output minute, rounded up.

import json, os, urllib.request

def api(path, body, key=None):
    headers = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"], "Content-Type": "application/json"}
    if key:
        headers["Idempotency-Key"] = key
    req = urllib.request.Request("https://api.sume.com" + path, json.dumps(body).encode(), headers)
    with urllib.request.urlopen(req) as r:
        body = json.load(r)
    return body.get("data", body)

wav = open("voiceover.wav", "rb").read()
seconds = int(os.environ["VOICEOVER_SECONDS"])
up = api("/v1/assets/upload-url", {"content_type": "audio/wav", "media_type": "audio",
                                   "size_bytes": len(wav), "filename": "voiceover.wav"})["upload"]
put = urllib.request.Request(up["url"], wav, up["headers"], method=up["method"])
urllib.request.urlopen(put).close()
asset = api(f"/v1/assets/{up['asset_id']}/complete", {})["asset"]
print(asset["status"], asset["url"])
plan = api("/v1/timeline-1.0/plan", {
    "audio": {"url": asset["url"], "duration_seconds": seconds},
    "video": [{"source_url": os.environ["CLIP_URL"], "start": 0, "duration": seconds}],
})
print(plan)

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume