Voice agent hits voicemail: write and record the message with TTS

ElevenLabs agents added Twilio answering-machine detection on 2026-09-07. What your agent should say to a machine, and how to pre-record it with Sume TTS.

4 min readSume
All posts

When an outbound voice agent detects an answering machine, it should leave one short, fixed message instead of improvising, and that message is best recorded ahead of time. ElevenLabs added Twilio answering-machine detection to its agent platform on 2026-09-07, per a changelog roundup, so the detection step now exists; the content of what to say is up to you. Sume does not run phone calls or detect machines, but its TTS can produce the recording in a telephony-friendly format.

This page covers the script, the format and a short checklist.

What should the voicemail message contain?

Keep it under 20 seconds: who is calling, why, one call-back number said twice, and nothing that depends on the listener's answers. Avoid personal data and anything a reviewer would read as a promise. Write numbers as words with spaces between digit groups so the voice reads them in pairs.

Because the text is fixed, you can render it once and reuse the file for every machine hit, which also removes a latency step at the moment the beep happens.

Which audio format works for a phone line?

Sume's TTS can return wav or raw with pcm_mulaw or pcm_alaw encoding at 8000 Hz, which are the encodings telephony stacks commonly take; the default is MP3 at 44.1 kHz. Confirm what your telephony provider accepts before you pick, because this page reads only Sume's side.

Sume TTS output options from the OpenAPI schema, and the reported ElevenLabs change, read 2026-10-02.
ChoiceValues Sume acceptsUse for the voicemail clip
Containermp3, wav, rawwav or raw for phone paths; mp3 for playback
Sample rate8000, 16000, 22050, 24000, 44100, 48000 Hz8000 Hz for narrowband lines
Encoding (wav/raw)pcm_f32le, pcm_s16le, pcm_mulaw, pcm_alawpcm_mulaw or pcm_alaw if your provider asks
Machine detectionNot a Sume featureElevenLabs reports Twilio detection from 2026-09-07

How do I render and store it?

Submit the text once with your chosen output_format, keep the job id, and download the audio from the result. Use an Idempotency-Key so a retry does not bill twice. When the job completes, Sume can notify you with a signed webhook, and you should keep polling as a backup, as the webhooks docs say.

Sume's TTS caps synthesized audio at 1,200 seconds per job, far above a voicemail clip. Details are in the API reference and Jobs and results.

What should I do?

Write three versions of the message, render them, play each on a real phone and keep the one people understand on the first listen. Leave the machine-detection logic with your agent platform and your telephony provider.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume