Can a MAI-Voice-2.1 WAV join a Sume timeline? Hosted audio only
Sume timeline audio joins only Sume-hosted files, and the import route takes TikTok or Instagram URLs, so a MAI-Voice-2.1 file stays outside. What to do.

No. A Sume timeline joins only audio that already lives on the Sume media host, so a WAV that MAI-Voice-2.1 produced in Microsoft's service cannot be dropped in as a part. Sume's Timeline audio page says every URL must already be this workspace's media.sume.com audio. If you want one gapless voiceover file on Sume, make the voice on Sume with TTS and join the takes there.
Microsoft announced MAI-Voice-2.1 at $22 per 1M characters and MAI-Voice-2.1-Flash at $15 per 1M characters on 2026-10-01 (Microsoft AI, read 2026-10-05). The question here is not price. It is where the finished file can go next.
What the docs say about hosting
The Timeline audio page tells you to import files first with POST /v1/media-imports. The request schema for that route in Sume's SDK types takes a public HTTPS TikTok or Instagram video or reel URL, mirrors the bytes into media.sume.com, and rejects YouTube and other hosts with unsupported_platform. The schema has no field for a loose WAV or MP3 file. So the import route does not turn your own audio into a Sume-hosted file.
The refusal you would meet on the join is stable and documented: unsupported_media_source when the URL is off-host, and source_not_found for a dead URL or one from a different workspace. Sume rejects off-host URLs at admit, before any join work starts.
Which input takes which kind of audio URL
Each Sume surface asks for a different kind of URL. Reading them side by side shows which one a MAI file could ever reach.
| Surface | Audio field | What it accepts |
|---|---|---|
| Timeline audio (concat, split) | parts[].url / url | This workspace's media.sume.com audio only |
| Timeline 1.0 render | audio.url or audio.parts[] | One Sume-hosted spine, or up to 20 gapless slices |
| Speech-to-text | audio_url | Public HTTPS URL, Sume media preferred |
| Video captions | video_url | Public HTTPS video that Sume can fetch |
Why the join step is the one that breaks
Speech-to-text and video captions read public HTTPS URLs, so a MAI file you host yourself can be transcribed. The two joining surfaces cannot read it. That split matters for a Flash workflow, because Flash makes up to 45 seconds of audio per call, and a five-minute script means several files that then need joining.
A timeline can join up to 20 parts, with a sample-domain join, no re-synthesis and no silence at the seams. That is the feature you lose when the parts are not Sume-hosted.
Two ways forward
If the goal is a joined voiceover inside a Sume video, generate the lines with Sume TTS. The result carries a Sume-hosted audio_url, so it can go straight into parts[]. At the list rate Sume TTS costs $0.0475 per 1,000 characters, against $0.015 for Flash and $0.022 for MAI-Voice-2.1 at Microsoft's per-million prices. A 1,000-character line is $0.0475 on Sume.
If the goal is MAI's voice itself, keep the finished track outside the timeline: host it on your own HTTPS storage, transcribe it with Sume STT for word timings, or mux it in your editor. State that choice in your pipeline so nobody expects a later Sume join to work.
A one-request test
Before you build around a mixed pipeline, run one check: submit a two-part concat with one Sume-hosted part and one URL from your own storage. Expect unsupported_media_source, a refusal at admit, before any join work starts. That one failed request settles the design for your team.
The Timeline audio job itself is a flat $0.01 per job and uses only worker ffmpeg, so the join is cheap once every part is Sume-hosted. Confirm the live rate in GET /v1/catalog.
There is also a hand-off cost to count. A pipeline that mixes two voice sources needs one more step per file: download the MAI output, store it, name it, and track which takes belong to which script. With Sume TTS the job result already holds the hosted audio URL, the duration and, if you ask for them, word timings and sentence segments. Fewer places to lose a take is a real saving even when the per-character price is higher.
Sources
Related posts
More in Integrations
- Make webhook 429 at 300 requests per 10 seconds and Sume retries
Make returns 429 above 300 webhook requests per 10 s. Sume retries a non-2xx up to 10 times, 30 s apart, so a short burst clears; your own replay loop does not.
- Make's webhook queue holds 667 items per 10,000 credits: plan for Sume
A Make webhook queue holds 667 items per 10,000 licensed credits, capped at 10,000, and answers 400 when full. Sume retries 10 times, so size the queue first.
- Make 'Process data in order' with Sume job events: what stays ordered
Make's Process data in order runs one execution at a time. Sume sends only terminal job events with no ordering promise, so dedupe on job_id.
- One Make webhook for Sume job.completed and format.run.terminal events
Sume sends job.completed for model jobs and format.run.terminal for Format runs, with one signature scheme. Branch on event, then on outcome, and dedupe by id.
Written by Sume