Can another vendor's audio go into a Sume timeline? What docs allow

Sume timeline audio joins only media.sume.com audio from your workspace. STT and avatar soundtracks take public HTTPS URLs. Which routes accept outside audio.

5 min readSume
All posts

Not directly by URL: Sume's timeline audio job accepts only audio that already lives on media.sume.com in your workspace, so a file from another voice or music vendor must get onto Sume first. The speech-to-text route is looser: it takes any public HTTPS audio_url.

That matters this week, when new voice sets and models arrive from other vendors almost daily. The routes below are from Sume's Timeline audio page, the Media inputs guide and the API reference, checked 2026-10-10.

Which routes take outside audio?

The table lists what the docs and schema say for each audio-taking route. "Public HTTPS" means a URL Sume can fetch; localhost, private-network, signed or private URLs are rejected.

Audio input rules by route (Sume docs and schema, checked 2026-10-10)
RouteAudio inputOutside URL accepted
STT 1.0 /v1/stt-1.0/transcribeaudio_urlYes, public HTTPS (Sume media URL preferred)
Timeline audio concat or splitparts[].url or urlNo, must be this workspace's media.sume.com audio
Audio detachvideo_url on the Sume media hostNo
Avatar Video package soundtrackpackage.soundtrack.audio_urlYes, public HTTPS, mirrored by Sume
Lip-sync routesSume-hosted audioCheck the route's field before relying on it

What does import cover?

POST /v1/media-imports takes a public HTTPS TikTok or Instagram video or reel URL; YouTube and other hosts are rejected with unsupported_platform. So it is for short-form video, not for a wav from a voice vendor. The media inputs guide adds that signed upload and download URLs are not part of the launch public API contract.

  • Voice or music file from another vendor: not importable through media-imports.
  • TikTok or Instagram reel: importable.
  • Sume's own outputs: already on media.sume.com.

What are the practical paths?

First, do the work on Sume. Generate narration with TTS and a track with Music; both land on Sume media URLs, so timeline audio can join them with no import. Second, if you need a transcript of outside audio, give STT a public HTTPS URL. Third, in the hosted MCP, the assets flow (create an upload URL, have the client PUT the bytes, then call assets_complete) exists for files, so ask whether your client and plan expose it before designing around it.

What if you must use an outside voice?

Keep that step outside Sume: produce the finished mix in the other tool, then caption the final video on Sume by giving the captions route a public HTTPS video_url. Video captions bill $0.20 per job for clips up to 60 seconds. You lose Sume's sample-exact joins for that audio, so decide the final level and length before you leave the other tool.

A decision guide

Ask three questions. Is the audio already made on Sume? Then use it directly in timeline audio, no import needed. Is it a file from another tool that you only need transcribed? Then host it at a public HTTPS URL and call STT, which costs $0.01 per audio minute and returns word timings. Is it a finished mix that should sit under a video? Then render the video elsewhere or caption the finished video, and keep Sume for the visual and caption steps.

Whichever path you take, record which tool made which audio. When the file source is mixed, that record is the only way to answer later where a given sound came from.

Which path fits (Sume rates and rules, checked 2026-10-10)
SituationPathCost note
Audio made on SumeTimeline audio directly$0.01 per join or split
Outside audio, need a transcriptSTT with a public HTTPS URL$0.01 per audio minute
Finished video with outside audioVideo captions on the video URL$0.20 per job, up to 60 s

Sources

Related posts

More in Integrations

All Integrations posts

Written by Sume