Decagon's Oct 1 launches: which of the four matter to a video team?
Voice 3, Personal Agent Gateway, Agent Modules and Duet Apprentice launched together. Only one touches voice; here is what each does and who needs it.

Decagon announced four products at Dialogues on October 1, 2026, and only Voice 3 (with its Chord speech model) is about audio. The other three concern agent traffic, industry modules and learning from escalated conversations, so a team making videos can ignore them.
The four, in Decagon's words
Decagon's Dialogues 2026 post describes them this way.
| Release | What Decagon says it does | Touches media? |
|---|---|---|
| Voice 3 | Duplex architecture with the Chord model for customer conversations | Yes, voice |
| Personal Agent Gateway | Identifies customers' personal agents and gives them a dedicated channel with permissions | No |
| Agent Modules | Extends the agent across journeys beyond support, for industries such as retail and telecom | No |
| Duet Apprentice | Learns from escalated conversations where a tenured rep handled an unusual case | No |
Voice 3 is the relevant one
Voice 3 is the part with sound. Decagon says the agent processes audio while speaking and narrates its progress during long tasks. That is a live-call product, and it does not replace a voiceover step in a video workflow.
What a video team still needs
Producing a video with narration is a file pipeline: write the script, generate speech, join the audio, add captions, render. Sume covers that pipeline through async jobs. tts_create creates the narration, timeline_audio joins parts, and a caption job burns the text. Rates come from the public rate card: text to speech at $0.0475 per 1,000 characters, a caption job at $0.20 for clips up to 60 seconds.
If your company also runs a Decagon agent, the two sit side by side. The agent answers calls, and your video pipeline makes the explainer clips that deflect some of them.
A quick triage
Do you need a phone or chat agent that talks with customers? Then read the Voice 3 page. Do you need media files? Then you need jobs. Do you handle personal-agent traffic or industry journeys? Read the other releases; they are not media tools.
Sources
Related posts
More in Comparisons
- Decagon Voice 3 handles 70+ languages mid-call. A TTS job takes one
Voice 3 detects language and switches mid-sentence. A Sume TTS job speaks one language per request. How to plan a multilingual script around that.
- Does Sume have a real-time avatar API? No, here is what it has instead
Sume has no live avatar session. It has async avatar jobs: create an avatar, render a talking video, or lip-sync a still to audio. Routes, limits and prices.
- Edits on desktop or an API render: which for a weekly Reel?
Edits now has a desktop app and an AI assistant. Use it for creative one-offs and a Sume render for repeated, logged Reels: a decision table with job prices.
- Eleven v4 Turbo in ElevenAgents vs Sume TTS as async jobs
Eleven v4 Turbo targets live agents. Sume TTS is an async job with poll or webhook and no streaming. Which one fits a call bot and which fits produced audio.
Written by Sume