Test call audio for voice agents: Sume TTS at 8 kHz mu-law
Generate repeatable phone-quality test utterances for a voice agent with Sume TTS output_format: 8000 Hz, pcm_mulaw. Fields, limits and a runnable script.

Voice agents get tested with the same handful of sentences over and over, so make those sentences as files. Sume's TTS endpoints accept output_format with a container, sample rate and encoding, which means you can ask for telephone-shaped audio: 8,000 Hz, pcm_mulaw. Pin the model and the voice, generate once per test case, and store the files.
The fields
| Field | Allowed values |
|---|---|
| output_format.container | mp3, wav, raw |
| output_format.sample_rate | 8000, 16000, 22050, 24000, 44100, 48000 |
| output_format.encoding | pcm_f32le, pcm_s16le, pcm_mulaw, pcm_alaw |
| transcript | up to 20,000 characters |
| voice.id | UUID or voi_ plus 32 hex characters |
Script
The script submits a job in async mode and prints the id. Fetching the result follows jobs and results: poll status, then call the result URL once the job is completed. A result requested early returns 409 job_not_completed.
import asyncio, os, httpx
CASES = {"refund": "I want to cancel my order.",
"spell": "My email is j, a, n, e at example dot com."}
async def main():
h = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
async with httpx.AsyncClient(headers=h, timeout=30) as c:
for name, text in CASES.items():
r = await c.post("https://api.sume.com/v1/tts-router/generate",
headers={"Idempotency-Key": "voice-qa-" + name}, json={
"model": "sonic-3.6", "transcript": text, "language": "en",
"voice": {"id": os.environ["SUME_VOICE_ID"]},
"output_format": {"container": "wav", "sample_rate": 8000,
"encoding": "pcm_mulaw"},
"mode": "async"})
r.raise_for_status()
print(name, r.json()["data"]["job"]["id"])
asyncio.run(main())Limits
- Narrowband audio hides detail that wideband hides less. A test that passes at 8 kHz may sound different on a 48 kHz web widget, so keep a wideband copy too.
- Timeline 1.0 audio output is wav or mp3, so do not expect it to emit mu-law. Keep mu-law turns as separate files, or convert on your side.
- An invalid voice id returns 400
invalid_voice_id; the voice must be a UUID orvoi_followed by 32 hex characters. - Each job is billed at $0.0475 per 1,000 characters (plus a 5.5% agent fee by default), so 100 short test lines cost cents. Repeated runs are separate paid jobs.
- Reference: API reference.
Related posts
More in Developers
- TTS voice.id: a UUID or a voi_ library id? What Sume accepts
Sume TTS voice.id takes a voice UUID or a voi_ library id. Any other shape fails with 400 invalid_voice_id before a job is queued or credits are reserved.
- TTS word timestamps to Timeline slide starts in Python
Call the Sume TTS Router with word timestamps, find the word that opens each slide, and build the Timeline video array of start times in a short Python script.
- TypeScript 7: an exhaustive switch over Sume job statuses
Turn a Sume job record into done, failed, canceled or running with a never check, so a new status breaks the build. Compiled with tsc 7.0.2 in strict mode.
- Bulk run says completed but UGC variants failed: read counts
A Sume bulk queue is completed once every item is terminal, not once every item succeeds. Read counts.failed and each item's status before shipping.
Written by Sume