Voice agent hits voicemail: write and record the message with TTS
ElevenLabs agents added Twilio answering-machine detection on 2026-09-07. What your agent should say to a machine, and how to pre-record it with Sume TTS.

When an outbound voice agent detects an answering machine, it should leave one short, fixed message instead of improvising, and that message is best recorded ahead of time. ElevenLabs added Twilio answering-machine detection to its agent platform on 2026-09-07, per a changelog roundup, so the detection step now exists; the content of what to say is up to you. Sume does not run phone calls or detect machines, but its TTS can produce the recording in a telephony-friendly format.
This page covers the script, the format and a short checklist.
What should the voicemail message contain?
Keep it under 20 seconds: who is calling, why, one call-back number said twice, and nothing that depends on the listener's answers. Avoid personal data and anything a reviewer would read as a promise. Write numbers as words with spaces between digit groups so the voice reads them in pairs.
Because the text is fixed, you can render it once and reuse the file for every machine hit, which also removes a latency step at the moment the beep happens.
Which audio format works for a phone line?
Sume's TTS can return wav or raw with pcm_mulaw or pcm_alaw encoding at 8000 Hz, which are the encodings telephony stacks commonly take; the default is MP3 at 44.1 kHz. Confirm what your telephony provider accepts before you pick, because this page reads only Sume's side.
| Choice | Values Sume accepts | Use for the voicemail clip |
|---|---|---|
| Container | mp3, wav, raw | wav or raw for phone paths; mp3 for playback |
| Sample rate | 8000, 16000, 22050, 24000, 44100, 48000 Hz | 8000 Hz for narrowband lines |
| Encoding (wav/raw) | pcm_f32le, pcm_s16le, pcm_mulaw, pcm_alaw | pcm_mulaw or pcm_alaw if your provider asks |
| Machine detection | Not a Sume feature | ElevenLabs reports Twilio detection from 2026-09-07 |
How do I render and store it?
Submit the text once with your chosen output_format, keep the job id, and download the audio from the result. Use an Idempotency-Key so a retry does not bill twice. When the job completes, Sume can notify you with a signed webhook, and you should keep polling as a backup, as the webhooks docs say.
Sume's TTS caps synthesized audio at 1,200 seconds per job, far above a voicemail clip. Details are in the API reference and Jobs and results.
What should I do?
Write three versions of the message, render them, play each on a real phone and keep the one people understand on the first listen. Leave the machine-detection logic with your agent platform and your telephony provider.
Sources
Related posts
More in Use cases
- Apartments.com photos: 2,048 px, JPG or PNG, no B&W
Apartments.com suggests 2,048 px on the longest side, JPG or PNG, owned by you and no black and white. Its two help pages disagree on GIF and file size.
- Apple Podcasts audio specs vs Sume TTS mp3 44.1 kHz 128 kbps
Apple Podcasts wants 44.1 kHz audio, about -16 LKFS and, for WAV or FLAC, stereo. What Sume TTS mp3 output meets and what you must still measure yourself.
- Apple Podcasts host-read video ads: cut the spot with Sume trim
Apple Podcasts lets creators insert video ads including host-read spots. Cut a host-read spot from an episode with Sume's exact-precision video trim.
- AI Act marking exemption for B2B and industrial output: how narrow
The Commission FAQ says a narrow Article 50(2) marking exemption is envisaged for B2B or industrial outputs, with conditions in the guidelines. What it lists.
Written by Sume