What is Decagon Chord? A speech model for support calls, not ad reads

Chord is Decagon Labs' voice model behind Voice 3, launched Oct 1, 2026. What the announcement says, and when you need a TTS job instead.

5 min readSume
All posts

Chord is the first voice model from Decagon Labs, built for customer conversations and used as the voice of Decagon Voice 3, which Decagon announced on October 1, 2026. It is a component of a customer-support agent product, so if you want a narrator for a video or a batch of ad reads, you want a text-to-speech job, not Chord.

The page we read describes Chord and Voice 3 as one launch. It lists no standalone Chord API, per-character price or voice catalog for narration, so treat any claim that you can call Chord like a TTS endpoint as unconfirmed.

What Decagon says Chord is

Decagon's announcement gives these facts about the model and the product around it.

  • Chord is the first voice model from Decagon Labs, trained for customer conversations.
  • It was post-trained on real-world customer experience conversations, on top of licensed data and consented voice talent.
  • It adjusts pace phrase by phrase, for example slowing down for a confirmation code or phone number, then returning to a conversational pace.
  • Voice 3 runs it on a duplex architecture, so the agent can process incoming audio while it speaks.
  • Voice 3 supports 70+ languages with automatic language detection.

Chord compared with a narration model

The difference is the job. A support voice reacts to a live caller. A narration model turns a finished script into a file you can cut into a video.

Which tool fits which job (read 2026-10-07)
JobFitsWhy
Answer a customer call, handle interruptionsDecagon Voice 3 (Chord)Duplex, live, built for support conversations
Voiceover for a 30-second adA TTS jobYou send a finished script and get an audio file
Transcript of a recorded callA speech-to-text jobAudio in, text and word timings out
Captions burned onto a clipA caption jobVideo in, captioned video out

What Sume covers

Sume covers the file-based side. The hosted MCP tools include tts_create and stt_create, and a TTS job is asynchronous: you submit it, then read the result from the job. The result carries the audio URL, duration, word timings and sentence slices.

Sume's public rate card lists text to speech at $0.0475 per 1,000 characters and transcription at $0.01 per audio minute. Check the live figure in GET /v1/catalog before you budget, because rates can change.

Sume does not run live phone calls. For that, pair a support-agent product such as Voice 3 with Sume for the recorded assets around it: the hold-music bed, the explainer video, the captions.

Decision shortcut

Ask whether a person is waiting on the other end in real time. If yes, you need a duplex agent. If you are producing a file, you need a job. Most video teams are in the second group.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume