Voice files from Gemini or ElevenLabs in a Sume timeline: 31 cents
Make the voice elsewhere, import the files into Sume media, join up to 20 parts for $0.01, render 3 minutes for $0.30. Total $0.31 plus your clips.

If you make the voice with another service, import the files into Sume media, join up to 20 of them into one file with a timeline audio concat ($0.01), and put that file under a Timeline 1.0 render as audio.url. A 3-minute video then costs $0.01 + 3 x $0.10 = $0.31 for the join and the render, before the cost of any clips you generate.
Sume does not list Gemini 3.8 TTS or ElevenLabs models, so this is an import path, not a model route (Sume API catalog, read 2026-10-09). The steps and limits are from the timeline and timeline audio docs.
The steps
Import each voice file with POST /v1/media-imports, since timeline audio only accepts Sume-hosted audio from your workspace and rejects off-host URLs at admit. Then call POST /v1/timeline-1.0/audio with operation: "concat" and parts[] of 1 to 20 items. Each part is { url, source_in?, duration? }, which lets you trim a part's head or tail without another job.
The result is kind: timeline_audio with an audio_url, a duration_seconds and segments[] giving each part's start offset. Use those offsets to set video[].start in the render. All parts must share a channel layout, or the job fails with audio_parts_channel_mismatch.
| Step | Endpoint | Cost |
|---|---|---|
| Import six files | POST /v1/media-imports | Not priced in this post |
| Concat six parts | POST /v1/timeline-1.0/audio | $0.01 |
| Render 180 seconds | POST /v1/timeline-1.0/render | 3 x $0.10 = $0.30 |
| Join plus render | $0.31 |
Formats and gaps
Keep wav where you have the choice: the docs say wav is sample-exact, and mp3 adds priming padding at every edge. If the vendor only gives you mp3, join once and avoid re-cutting it. The join is sample-domain with no silence added at the seams, so insert your own pauses as part of the script or with a silent spine segment.
Set audio.duration_seconds to the joined length (1 to 1,800). The render fails with audio_parts_shorter_than_duration if you declare more than the parts provide.
Check the license first
Each voice service has its own terms for commercial use and for derived video. Read the vendor's page before you ship ads with its audio; Sume's tools process the file but do not change those terms. For pricing context see the related notes on per-minute speech cost.
When you do not need the concat
If the voice is one file already, skip the join and set audio.url to the imported file. The render alone is then $0.10 per output minute, so a 3-minute video is $0.30. Use the concat only when you have several takes, and split takes at 20 per job.
If a take has a bad second in the middle, split the file with a timeline audio split and re-join the good ranges: that is two jobs at $0.01 each.
Poll each job by id and use an Idempotency-Key on the concat and the render, since a retry with the same key returns the original job and does not bill twice.
Sources
Related posts
More in Integrations
- VS Code MCP: Reset Trust after you grant Sume the Write scope
Write on Sume's consent page changes the server's abilities, not VS Code's trust decision. MCP: Reset Trust makes VS Code ask again; mcp_health shows scopes.
- VS Code Settings Sync and MCP servers: keep the Sume key out
VS Code can sync MCP config when the MCP Servers sync option is on. Keep the Sume API key out of the file with an input variable, or use OAuth instead.
- VS Code user MCP config: a read-only and a Write Sume profile
VS Code's MCP: Open User Configuration edits per-user servers. Keep OAuth read-only Sume in your everyday setup and an API key entry only where paid jobs run.
- VS Code 1.141 background shells and a long Sume render
VS Code 1.141 tracks background shells for agent sessions. Start a Sume render, keep the job id, and use sume jobs watch instead of repeating the paid create.
Written by Sume