What is Decagon Chord? A speech model for support calls, not ad reads
Chord is Decagon Labs' voice model behind Voice 3, launched Oct 1, 2026. What the announcement says, and when you need a TTS job instead.

Chord is the first voice model from Decagon Labs, built for customer conversations and used as the voice of Decagon Voice 3, which Decagon announced on October 1, 2026. It is a component of a customer-support agent product, so if you want a narrator for a video or a batch of ad reads, you want a text-to-speech job, not Chord.
The page we read describes Chord and Voice 3 as one launch. It lists no standalone Chord API, per-character price or voice catalog for narration, so treat any claim that you can call Chord like a TTS endpoint as unconfirmed.
What Decagon says Chord is
Decagon's announcement gives these facts about the model and the product around it.
- Chord is the first voice model from Decagon Labs, trained for customer conversations.
- It was post-trained on real-world customer experience conversations, on top of licensed data and consented voice talent.
- It adjusts pace phrase by phrase, for example slowing down for a confirmation code or phone number, then returning to a conversational pace.
- Voice 3 runs it on a duplex architecture, so the agent can process incoming audio while it speaks.
- Voice 3 supports 70+ languages with automatic language detection.
Chord compared with a narration model
The difference is the job. A support voice reacts to a live caller. A narration model turns a finished script into a file you can cut into a video.
| Job | Fits | Why |
|---|---|---|
| Answer a customer call, handle interruptions | Decagon Voice 3 (Chord) | Duplex, live, built for support conversations |
| Voiceover for a 30-second ad | A TTS job | You send a finished script and get an audio file |
| Transcript of a recorded call | A speech-to-text job | Audio in, text and word timings out |
| Captions burned onto a clip | A caption job | Video in, captioned video out |
What Sume covers
Sume covers the file-based side. The hosted MCP tools include tts_create and stt_create, and a TTS job is asynchronous: you submit it, then read the result from the job. The result carries the audio URL, duration, word timings and sentence slices.
Sume's public rate card lists text to speech at $0.0475 per 1,000 characters and transcription at $0.01 per audio minute. Check the live figure in GET /v1/catalog before you budget, because rates can change.
Sume does not run live phone calls. For that, pair a support-agent product such as Voice 3 with Sume for the recorded assets around it: the hold-music bed, the explainer video, the captions.
Decision shortcut
Ask whether a person is waiting on the other end in real time. If yes, you need a duplex agent. If you are producing a file, you need a job. Most video teams are in the second group.
Sources
Related posts
More in Comparisons
- Can I use Suno Speech beta for an ad voiceover? What to check first
Suno Speech beta makes one track with voice and music. Before using it for ads, check price, languages and edits, then see how Sume splits voice and music.
- Which TikTok ad placements take a 10-minute video? Only auction
TikTok's auction non-Spark page allows up to 10 minutes; reservation in-feed and TopView stop at 60 seconds. Pick the placement first, then cut with video-trim.
- Zoom in on a small product in frame: AI recompose or crop and upscale
Product too small in the photo? A Pillow crop plus Sume Image Upscale ($0.20) keeps real pixels; an Ideogram 4.5 recompose edit costs $0.075 and redraws them.
- Sume vs Argil: AI avatar video and video agents compared
Argil makes AI-avatar and story videos with a chat agent, Director; Sume is a video agent with a multi-model API. Avatars, API, pricing, and limits compared.
Written by Sume