Speechify API vs Sume: speech marks and streaming vs video jobs

Speechify's API streams speech with word-level marks and voice cloning. Sume documents no standalone speech endpoint, but accepts audio for talking-video jobs.

5 min readSume
All posts

Is Speechify's API an alternative to anything in Sume?

They barely overlap. Speechify's API is text-to-speech: one endpoint, two models, streaming, voice cloning and word timestamps. Sume's docs describe no standalone text-to-speech endpoint. The two meet only if you generate speech with Speechify and then give the audio to a Sume talking-clip route. For voice work alone, Speechify is the better and more direct fit.

What does the Speechify API document?

The core endpoint is POST https://api.speechify.ai/v1/audio/speech. Two models are active: simba-3.2, described as the lowest time to first byte and richest expressivity and the recommended Simba 3 model for English, and simba-3.0, the default, which supports English, German, Spanish, French, Italian and Brazilian Portuguese. The retired simba-english and simba-multilingual models return 400 model_retired.

Features listed: streaming so playback starts before the full audio is generated (up to 20,000 characters per request), voice cloning from 10 to 30 second samples with speaker consent, SSML control of pitch, rate, pauses, emphasis and 13 emotion presets, and speech marks that give word-level timestamps for sync and captions. One Bearer key covers every endpoint, and there are plugins for LiveKit, Pipecat, Vapi and Deepgram.

Voice features against Sume's public surface (read 2026-10-02)
FeatureSpeechify APISume
Text to speech endpointPOST /v1/audio/speechNo standalone endpoint in the docs
Streaming playbackYes, up to 20,000 charactersNot documented; Sume jobs are async
Voice cloningFrom 10 to 30 s samples with consentNot documented as a voice API
Word timestampsSpeech marksCaption jobs align a script, not a speech-marks API
Video from a scriptNot what it is foravatar-1.0/talking-video, 4 to 60 s

Where does Sume pick up after the audio exists?

The Fabric route, POST /v1/veed/fabric-1.0, takes an audio_url, a measured duration_seconds and one visual source (image_url or avatar_handle, never both). That lets you keep Speechify's voice and have Sume render the talking still. Because Sume's docs say video models do not lip-sync to generated speech laid underneath later, a talking face goes through Fabric with the still and audio together rather than a video clip with narration added afterwards.

For captions, POST /v1/video-captions burns styled captions onto a public video URL and can take a language hint and a script for alignment. It does not expose word timestamps as a separate product, so if you need the marks themselves for your own player, take them from Speechify.

What are the limits to plan around?

On Speechify, the 20,000 character streaming cap and the sample length for cloning are the numbers to design around. On Sume, plan concurrency is the first limit you meet: Free 1, Pro 4, Startup 8, Scale 20 jobs processing, with a queue behind it. Webhooks are terminal-only and signed, so keep a poll fallback. The avatar-video window of 4 to 60 seconds means a long narration must be split across jobs.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: spch-cap-001" \
  -d '{"video_url":"https://example.com/clip.mp4","language":"auto"}'

How do the two handle retries and errors?

Speechify's overview gives one clear error example: retired models return 400 model_retired, so pin a current model such as simba-3.2 or simba-3.0 and handle that code if you cache model ids. I did not read its full error list.

Sume's job-based calls have a full public error table: 400 invalid_request, 401 unauthorized, 402 insufficient_credits, 409 conflicts, 413 payload_too_large, 429 rate_limited and 429 queue_full, and 503 for provider capacity. Failed jobs add a category and a retryability flag. Pair every paid submit with an Idempotency-Key so a retried request returns the original job, and never resubmit just because a local timeout fired.

Which should you use?

Use Speechify for the voice itself, especially when you need streaming and cloning. Use Sume for the finished clip. If you only need captions on an existing video, the Sume caption route is a smaller step than a full avatar job. Read the caption guide for style names and language handling.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume