Speechify API vs Sume: speech marks and streaming vs video jobs
Speechify's API streams speech with word-level marks and voice cloning. Sume documents no standalone speech endpoint, but accepts audio for talking-video jobs.

Is Speechify's API an alternative to anything in Sume?
They barely overlap. Speechify's API is text-to-speech: one endpoint, two models, streaming, voice cloning and word timestamps. Sume's docs describe no standalone text-to-speech endpoint. The two meet only if you generate speech with Speechify and then give the audio to a Sume talking-clip route. For voice work alone, Speechify is the better and more direct fit.
What does the Speechify API document?
The core endpoint is POST https://api.speechify.ai/v1/audio/speech. Two models are active: simba-3.2, described as the lowest time to first byte and richest expressivity and the recommended Simba 3 model for English, and simba-3.0, the default, which supports English, German, Spanish, French, Italian and Brazilian Portuguese. The retired simba-english and simba-multilingual models return 400 model_retired.
Features listed: streaming so playback starts before the full audio is generated (up to 20,000 characters per request), voice cloning from 10 to 30 second samples with speaker consent, SSML control of pitch, rate, pauses, emphasis and 13 emotion presets, and speech marks that give word-level timestamps for sync and captions. One Bearer key covers every endpoint, and there are plugins for LiveKit, Pipecat, Vapi and Deepgram.
| Feature | Speechify API | Sume |
|---|---|---|
| Text to speech endpoint | POST /v1/audio/speech | No standalone endpoint in the docs |
| Streaming playback | Yes, up to 20,000 characters | Not documented; Sume jobs are async |
| Voice cloning | From 10 to 30 s samples with consent | Not documented as a voice API |
| Word timestamps | Speech marks | Caption jobs align a script, not a speech-marks API |
| Video from a script | Not what it is for | avatar-1.0/talking-video, 4 to 60 s |
Where does Sume pick up after the audio exists?
The Fabric route, POST /v1/veed/fabric-1.0, takes an audio_url, a measured duration_seconds and one visual source (image_url or avatar_handle, never both). That lets you keep Speechify's voice and have Sume render the talking still. Because Sume's docs say video models do not lip-sync to generated speech laid underneath later, a talking face goes through Fabric with the still and audio together rather than a video clip with narration added afterwards.
For captions, POST /v1/video-captions burns styled captions onto a public video URL and can take a language hint and a script for alignment. It does not expose word timestamps as a separate product, so if you need the marks themselves for your own player, take them from Speechify.
What are the limits to plan around?
On Speechify, the 20,000 character streaming cap and the sample length for cloning are the numbers to design around. On Sume, plan concurrency is the first limit you meet: Free 1, Pro 4, Startup 8, Scale 20 jobs processing, with a queue behind it. Webhooks are terminal-only and signed, so keep a poll fallback. The avatar-video window of 4 to 60 seconds means a long narration must be split across jobs.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: spch-cap-001" \
-d '{"video_url":"https://example.com/clip.mp4","language":"auto"}'How do the two handle retries and errors?
Speechify's overview gives one clear error example: retired models return 400 model_retired, so pin a current model such as simba-3.2 or simba-3.0 and handle that code if you cache model ids. I did not read its full error list.
Sume's job-based calls have a full public error table: 400 invalid_request, 401 unauthorized, 402 insufficient_credits, 409 conflicts, 413 payload_too_large, 429 rate_limited and 429 queue_full, and 503 for provider capacity. Failed jobs add a category and a retryability flag. Pair every paid submit with an Idempotency-Key so a retried request returns the original job, and never resubmit just because a local timeout fired.
Which should you use?
Use Speechify for the voice itself, especially when you need streaming and cloning. Use Sume for the finished clip. If you only need captions on an existing video, the Sume caption route is a smaller step than a full avatar job. Read the caption guide for style names and language handling.
Sources
Related posts
More in Comparisons
- StepAudio 3 Gen: voice, SFX and music in one clip vs Sume jobs
StepFun's stepaudio-3-gen-preview makes voice, effects, ambience and music in one audio output. Sume uses separate music, speech and Timeline mix steps.
- StepAudio 3 Music from a dry vocal or reference audio vs Sume
StepAudio 3 Music accepts lyrics, vocals or reference audio, even scoring a dry vocal. Sume Music takes a text prompt and one optional image, no audio.
- Synthesia burned-in captions on dubbed videos, and Sume captions
Synthesia added one-click burned-in captions to dubbed videos on 9/30/2026. Here is what that means, and how to burn captions onto a finished video URL on Sume.
- Which Synthesia plan includes API access? Pro is limited
Synthesia lists no API on Basic or Starter, a limited API with 360 minutes a year on Pro, and full access on Enterprise. What a Sume API key gives you instead.
Written by Sume