Voice files from Gemini or ElevenLabs in a Sume timeline: 31 cents

Make the voice elsewhere, import the files into Sume media, join up to 20 parts for $0.01, render 3 minutes for $0.30. Total $0.31 plus your clips.

5 min readSume
All posts

If you make the voice with another service, import the files into Sume media, join up to 20 of them into one file with a timeline audio concat ($0.01), and put that file under a Timeline 1.0 render as audio.url. A 3-minute video then costs $0.01 + 3 x $0.10 = $0.31 for the join and the render, before the cost of any clips you generate.

Sume does not list Gemini 3.8 TTS or ElevenLabs models, so this is an import path, not a model route (Sume API catalog, read 2026-10-09). The steps and limits are from the timeline and timeline audio docs.

The steps

Import each voice file with POST /v1/media-imports, since timeline audio only accepts Sume-hosted audio from your workspace and rejects off-host URLs at admit. Then call POST /v1/timeline-1.0/audio with operation: "concat" and parts[] of 1 to 20 items. Each part is { url, source_in?, duration? }, which lets you trim a part's head or tail without another job.

The result is kind: timeline_audio with an audio_url, a duration_seconds and segments[] giving each part's start offset. Use those offsets to set video[].start in the render. All parts must share a channel layout, or the job fails with audio_parts_channel_mismatch.

Import-and-render budget for six voice files and a 3-minute video, as of 2026-10-09
StepEndpointCost
Import six filesPOST /v1/media-importsNot priced in this post
Concat six partsPOST /v1/timeline-1.0/audio$0.01
Render 180 secondsPOST /v1/timeline-1.0/render3 x $0.10 = $0.30
Join plus render$0.31

Formats and gaps

Keep wav where you have the choice: the docs say wav is sample-exact, and mp3 adds priming padding at every edge. If the vendor only gives you mp3, join once and avoid re-cutting it. The join is sample-domain with no silence added at the seams, so insert your own pauses as part of the script or with a silent spine segment.

Set audio.duration_seconds to the joined length (1 to 1,800). The render fails with audio_parts_shorter_than_duration if you declare more than the parts provide.

Check the license first

Each voice service has its own terms for commercial use and for derived video. Read the vendor's page before you ship ads with its audio; Sume's tools process the file but do not change those terms. For pricing context see the related notes on per-minute speech cost.

When you do not need the concat

If the voice is one file already, skip the join and set audio.url to the imported file. The render alone is then $0.10 per output minute, so a 3-minute video is $0.30. Use the concat only when you have several takes, and split takes at 20 per job.

If a take has a bad second in the middle, split the file with a timeline audio split and re-join the good ranges: that is two jobs at $0.01 each.

Poll each job by id and use an Idempotency-Key on the concat and the render, since a retry with the same key returns the original job and does not bill twice.

Sources

Related posts

More in Integrations

All Integrations posts

Written by Sume