Respeecher Space at $2 an hour vs Sume async text to speech
Respeecher Space is a real-time TTS API for voice agents at $2 an hour. Sume TTS is async and per character. Which fits narration, and which fits live voice.

Respeecher Space is a real-time text-to-speech API aimed at voice agents, priced at $2 an hour pay-as-you-go; Sume's text to speech is asynchronous and per character, so it suits narration, not a live conversation. Use Respeecher for a voice agent that must talk back in real time, and Sume for finished audio that goes into a video or timeline.
Respeecher facts are from its home page, read on 2026-10-02. Sume facts are from its OpenAPI reference and docs.
What does Respeecher list?
The page describes Respeecher Space as a real-time text-to-speech API for voice agents, pay as you go at $2 an hour, cancel anytime. It also lists a Voice Marketplace of 40+ AI voices with a plugin for Pro Tools, language-agnostic output, emotion transfer, and strict consent mechanisms with transparent voice origin tracking.
The page gives no latency figure, language list or character limit, so those are not compared.
Is Sume's text to speech real-time?
No. The OpenAPI description of POST /v1/tts-1.0/generate states that phase 1 is an async job with poll or webhook and is non-streaming. mode: async returns at once with status, result, events and cancel URLs; mode: sync and subscribe are aliases for a bounded wait of up to 30 seconds, and a job still running after that must be polled rather than resubmitted.
That fits narration, captions, lip-sync audio and voice-overs, where you send the whole script and receive a file. It does not fit a call-center agent that must start speaking in a few hundred milliseconds.
How do the two compare on cost?
Respeecher bills by time, Sume by characters. At an assumed 900 characters per minute (150 words per minute at 6 characters per word, an assumption), an hour of speech is about 54,000 characters. Multiply that by the per-character price that GET /v1/tts-router/models returns for the model you pick (the TTS Router list price times 1.10). The two are not the same unit, though: Respeecher's hour may be connection or generated time, which its page does not say, so compare on your own volume.
| Property | Respeecher Space (page) | Sume TTS (docs) |
|---|---|---|
| Delivery | Real-time, for voice agents | Async job, non-streaming |
| Pricing unit | $2 per hour, pay-as-you-go | Per character; price in GET /v1/tts-router/models |
| Hour of speech (about 54,000 characters, assumed) | $2 | Characters times the catalog per-character price |
| Voice cloning | Core product, with consent framework | No clone route in the API reference |
| Output for video work | Not described | MP3, WAV or raw; word timings; sentence segments |
What about voice cloning and consent?
Respeecher's page puts consent and voice tracking at the centre. Sume's API reference selects voices by id, by avatar id or by avatar handle, and rejects a voice name from another vendor with invalid_voice_id before a job is queued. There is no route to upload a sample and clone a voice in the OpenAPI paths we read, so if cloning a specific performer is the requirement, Sume is not the product.
Where Sume does help is the follow-on work. Word timings (timestamps: { words: true }) and sentence segmentation turn a voice track into caption cues and timeline slide starts; see Jobs and results for polling and signed webhooks.
curl -X POST https://api.sume.com/v1/tts-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: timings-001" \
-d '{
"model": "sonic-latest",
"transcript": "Here is the plan. First, we record. Then we edit.",
"avatar_handle": "@my.avatar",
"timestamps": { "words": true },
"mode": "sync",
"wait_timeout_seconds": 30
}'Which should I choose?
For a live voice agent, Sume's async jobs are the wrong tool and Respeecher's real-time API is the one described for it. For narration where a script is complete before synthesis, Sume's per-character jobs and timing data fit, and price both on your own volume, since an hour on Respeecher's page is not defined.
Sources
Related posts
More in Comparisons
- Runway lists 18 video model ids: which nine does Sume have?
Of the 18 video ids on Runway's models page, nine match a Sume id: four Seedance rows, MiniMax H3 and Max, Wan 3.0, Grok Imagine, Gemini Omni Flash 1.1.
- Segmind API vs Sume: PixelFlow workflows or saved Formats
Segmind turns visual PixelFlow graphs into API endpoints. Sume saves an agent thread as a Format you call over the API. How the two reuse recipes.
- Sonilo segment-level music controls vs Sume section markers
Sonilo's text-to-music lets you set styles and moods per section. On Sume you write section markers like [0:00-0:30] Intro: inside one 5000-character prompt.
- Sonilo video-to-music on fal.ai: 600 s of footage vs Sume's route
Sonilo's video-to-music model scores footage up to 600 seconds on fal.ai. Sume has no video-to-music call: inspect the clip, write a prompt, mix with Timeline.
Written by Sume