Test call audio for voice agents: Sume TTS at 8 kHz mu-law

Generate repeatable phone-quality test utterances for a voice agent with Sume TTS output_format: 8000 Hz, pcm_mulaw. Fields, limits and a runnable script.

4 min readSume
All posts

Voice agents get tested with the same handful of sentences over and over, so make those sentences as files. Sume's TTS endpoints accept output_format with a container, sample rate and encoding, which means you can ask for telephone-shaped audio: 8,000 Hz, pcm_mulaw. Pin the model and the voice, generate once per test case, and store the files.

The fields

Fields on POST /v1/tts-router/generate, from the Sume OpenAPI reference, read 2026-10-01.
FieldAllowed values
output_format.containermp3, wav, raw
output_format.sample_rate8000, 16000, 22050, 24000, 44100, 48000
output_format.encodingpcm_f32le, pcm_s16le, pcm_mulaw, pcm_alaw
transcriptup to 20,000 characters
voice.idUUID or voi_ plus 32 hex characters

Script

The script submits a job in async mode and prints the id. Fetching the result follows jobs and results: poll status, then call the result URL once the job is completed. A result requested early returns 409 job_not_completed.

import asyncio, os, httpx

CASES = {"refund": "I want to cancel my order.",
         "spell": "My email is j, a, n, e at example dot com."}

async def main():
    h = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
    async with httpx.AsyncClient(headers=h, timeout=30) as c:
        for name, text in CASES.items():
            r = await c.post("https://api.sume.com/v1/tts-router/generate",
                             headers={"Idempotency-Key": "voice-qa-" + name}, json={
                "model": "sonic-3.6", "transcript": text, "language": "en",
                "voice": {"id": os.environ["SUME_VOICE_ID"]},
                "output_format": {"container": "wav", "sample_rate": 8000,
                                  "encoding": "pcm_mulaw"},
                "mode": "async"})
            r.raise_for_status()
            print(name, r.json()["data"]["job"]["id"])

asyncio.run(main())

Limits

  • Narrowband audio hides detail that wideband hides less. A test that passes at 8 kHz may sound different on a 48 kHz web widget, so keep a wideband copy too.
  • Timeline 1.0 audio output is wav or mp3, so do not expect it to emit mu-law. Keep mu-law turns as separate files, or convert on your side.
  • An invalid voice id returns 400 invalid_voice_id; the voice must be a UUID or voi_ followed by 32 hex characters.
  • Each job is billed at $0.0475 per 1,000 characters (plus a 5.5% agent fee by default), so 100 short test lines cost cents. Repeated runs are separate paid jobs.
  • Reference: API reference.

Related posts

More in Developers

All Developers posts

Written by Sume