Can a MAI-Voice-2.1 WAV join a Sume timeline? Hosted audio only

Sume timeline audio joins only Sume-hosted files, and the import route takes TikTok or Instagram URLs, so a MAI-Voice-2.1 file stays outside. What to do.

4 min readSume
All posts

No. A Sume timeline joins only audio that already lives on the Sume media host, so a WAV that MAI-Voice-2.1 produced in Microsoft's service cannot be dropped in as a part. Sume's Timeline audio page says every URL must already be this workspace's media.sume.com audio. If you want one gapless voiceover file on Sume, make the voice on Sume with TTS and join the takes there.

Microsoft announced MAI-Voice-2.1 at $22 per 1M characters and MAI-Voice-2.1-Flash at $15 per 1M characters on 2026-10-01 (Microsoft AI, read 2026-10-05). The question here is not price. It is where the finished file can go next.

What the docs say about hosting

The Timeline audio page tells you to import files first with POST /v1/media-imports. The request schema for that route in Sume's SDK types takes a public HTTPS TikTok or Instagram video or reel URL, mirrors the bytes into media.sume.com, and rejects YouTube and other hosts with unsupported_platform. The schema has no field for a loose WAV or MP3 file. So the import route does not turn your own audio into a Sume-hosted file.

The refusal you would meet on the join is stable and documented: unsupported_media_source when the URL is off-host, and source_not_found for a dead URL or one from a different workspace. Sume rejects off-host URLs at admit, before any join work starts.

Which input takes which kind of audio URL

Each Sume surface asks for a different kind of URL. Reading them side by side shows which one a MAI file could ever reach.

Audio inputs by surface (read 2026-10-05)
SurfaceAudio fieldWhat it accepts
Timeline audio (concat, split)parts[].url / urlThis workspace's media.sume.com audio only
Timeline 1.0 renderaudio.url or audio.parts[]One Sume-hosted spine, or up to 20 gapless slices
Speech-to-textaudio_urlPublic HTTPS URL, Sume media preferred
Video captionsvideo_urlPublic HTTPS video that Sume can fetch

Why the join step is the one that breaks

Speech-to-text and video captions read public HTTPS URLs, so a MAI file you host yourself can be transcribed. The two joining surfaces cannot read it. That split matters for a Flash workflow, because Flash makes up to 45 seconds of audio per call, and a five-minute script means several files that then need joining.

A timeline can join up to 20 parts, with a sample-domain join, no re-synthesis and no silence at the seams. That is the feature you lose when the parts are not Sume-hosted.

Two ways forward

If the goal is a joined voiceover inside a Sume video, generate the lines with Sume TTS. The result carries a Sume-hosted audio_url, so it can go straight into parts[]. At the list rate Sume TTS costs $0.0475 per 1,000 characters, against $0.015 for Flash and $0.022 for MAI-Voice-2.1 at Microsoft's per-million prices. A 1,000-character line is $0.0475 on Sume.

If the goal is MAI's voice itself, keep the finished track outside the timeline: host it on your own HTTPS storage, transcribe it with Sume STT for word timings, or mux it in your editor. State that choice in your pipeline so nobody expects a later Sume join to work.

A one-request test

Before you build around a mixed pipeline, run one check: submit a two-part concat with one Sume-hosted part and one URL from your own storage. Expect unsupported_media_source, a refusal at admit, before any join work starts. That one failed request settles the design for your team.

The Timeline audio job itself is a flat $0.01 per job and uses only worker ffmpeg, so the join is cheap once every part is Sume-hosted. Confirm the live rate in GET /v1/catalog.

There is also a hand-off cost to count. A pipeline that mixes two voice sources needs one more step per file: download the MAI output, store it, name it, and track which takes belong to which script. With Sume TTS the job result already holds the hosted audio URL, the duration and, if you ask for them, word timings and sentence segments. Fewer places to lose a take is a real saving even when the per-character price is higher.

Sources

Related posts

More in Integrations

All Integrations posts

Written by Sume