Decagon Voice 3 duplex agent or a voiceover job: which do you need?
Voice 3 listens while it speaks. A voiceover job does not. How to choose between a live voice agent and a file-based TTS job for your project.

Pick a duplex voice agent like Decagon Voice 3 when a person talks back to the system in real time; pick a file-based voiceover job when you already know the words. Voice 3, announced October 1, 2026, can process incoming audio while it is speaking, which a script-to-audio job does not do.
What duplex means in Decagon's design
Decagon says Voice 3 removes the delays of a cascaded pipeline, the silent gaps and interruptions that come from passing audio between separate steps. It describes two layers: a low-latency conversational model that listens and speaks, and a more powerful model that handles reasoning, tool calling and guardrails.
That structure exists to survive a live conversation. The caller can interrupt, ask a follow-up, or change topic, and the agent keeps up.
What a voiceover job does instead
A voiceover job has no caller. You send the exact words, pick a voice, and wait for a file. In Sume, the hosted tool is tts_create. It needs an idempotency key, a transcript of 1 to 20,000 characters and a voice selector. The job returns an audio URL, duration, word timings and sentence segments you can feed into a timeline render.
The trade is control for reactivity. Because the script is fixed, you can check it, approve it and reuse unchanged takes. Nothing in the file changes depending on what a listener says.
Side by side
The table below separates the two.
| Question | Duplex voice agent | Voiceover job |
|---|---|---|
| Is a person talking back? | Yes, live | No |
| Who writes the words? | The agent, at runtime | You, before the job |
| Output | A conversation | An audio file with timings |
| Can you review it first? | Only by testing | Yes, read the script and listen |
| Typical use | Support, bookings, call centers | Ads, explainers, courses, shorts |
Cost shape
A live agent is priced per conversation or per minute of a service. A job is priced per character. Sume's rate card lists text to speech at $0.0475 per 1,000 characters, with a one-cent minimum per job and a 95-cent ceiling at the 20,000-character limit. Confirm the current rate in GET /v1/catalog.
If you only need a recorded greeting, a menu prompt or a hold message for a phone system, that is a job, not an agent. Generate it once, check it, and load the file.
When to use both
Some teams run a live agent for calls and use jobs for the assets around it: the intro video on the support page, the narrated how-to, the captioned clip in the help center. They share a brand voice only if you pick it deliberately in each tool.
Sources
Related posts
More in Comparisons
- Does Sume have a real-time avatar API? No, here is what it has instead
Sume has no live avatar session. It has async avatar jobs: create an avatar, render a talking video, or lip-sync a still to audio. Routes, limits and prices.
- Edits on desktop or an API render: which for a weekly Reel?
Edits now has a desktop app and an AI assistant. Use it for creative one-offs and a Sume render for repeated, logged Reels: a decision table with job prices.
- Eleven v4 Turbo in ElevenAgents vs Sume TTS as async jobs
Eleven v4 Turbo targets live agents. Sume TTS is an async job with poll or webhook and no streaming. Which one fits a call bot and which fits produced audio.
- ElevenLabs Flash $0.04 and v3 $0.08 vs Sume TTS, 20,000 characters
ElevenLabs lists $0.04 (Flash) and $0.08 (v3) per 1,000 characters; Sume TTS is $0.0475. A full 20,000-character Sume job is $0.95 before the fee.
Written by Sume