Murf API alternative for video voiceover: where Sume fits
Murf is a voice API with Falcon, dubbing and 150+ voices. Sume has an async TTS route and turns a script into a talking video; it has no dubbing API.
Is Sume an alternative to the Murf API?
Only partly. Sume's OpenAPI schema lists POST /v1/tts-1.0/generate, an async, non-streaming text-to-speech job that returns audio artifacts, but it has no dubbing or translation API and no real-time audio. Murf is a voice product with those. Sume's strength is turning a script into a finished talking-head video. If your deliverable is low-latency or dubbed audio, pick Murf. If it is a video with a presenter, read on.
What does the Murf API offer?
Murf's API page lists 150+ voices in 35 languages, including MultiNative voices that work across languages with the same voice. Its primary model is Falcon, quoted at 55 ms model latency and sub-130 ms time to first audio, with a second option called Gen 2 TTS that has customisable voice parameters. Pricing is quoted at 1 cent per minute with no tiers, and the page claims 10,000 concurrent calls at the same latency.
Beyond TTS it lists a Voice Changer API, a Dubbing API for 25+ languages, a Translation API for 21+ languages, pitch, speed and prosody control, custom pronunciation and 15+ expressive styles. A Python SDK is described as production-ready. None of that is replicated by Sume.
| Need | Murf API | Sume |
|---|---|---|
| Voice from text | Yes; Falcon and Gen 2 | Yes; POST /v1/tts-1.0/generate, async non-streaming job |
| Real-time streaming audio | Falcon, sub-130 ms to first audio (vendor figure) | No; TTS is an async job with polling or webhook, not streaming |
| Dub or translate audio | Dubbing and Translation APIs | Not offered as an API |
| Presenter video from a script | Not what the page describes | POST /v1/avatar-1.0/talking-video |
| Burned-in captions | Not covered | POST /v1/video-captions |
What does Sume do with a script?
With Avatar 1.0 you create a reusable avatar, then call POST /v1/avatar-1.0/talking-video with an avatar_handle and either a script or ordered video_inputs. Sume accepts scripts it estimates at 4 to 60 seconds. Quality is standard, plus (default) or max, aspect ratio defaults to 9:16, and resolution is currently 720p. Inline captions can be switched on in the same request.
For speech alone, the TTS route above returns an audio artifact you can fetch from the job result.
Can you use Murf audio with Sume video?
Yes, in one documented place. The VEED Fabric 1.0 route (POST /v1/veed/fabric-1.0) takes an audio_url, a measured duration_seconds and exactly one visual source: an image_url or an avatar_handle. So you can render speech elsewhere, host it at a public HTTPS URL, and have Sume animate a still to that audio. The models overview lists Fabric's older alias price at $0.1875 per second at 720p, and a MiniMax H3 Max lip sync route that takes the same still-plus-audio body with audio of 5 to 14.8 seconds.
Measure the audio duration yourself and send it as duration_seconds, rather than guessing.
curl -X POST https://api.sume.com/v1/veed/fabric-1.0 \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: murf-fabric-001" \
-d '{"image_url":"https://example.com/presenter.png","audio_url":"https://example.com/line.mp3","duration_seconds":8}'What does a script-to-video job cost in planning terms?
Murf quotes 1 cent per minute for voice. Sume's price for a talking video is not a per-minute voice rate, because a job includes the avatar render. The docs show the retiring Fabric alias at $0.1875 per second at 720p, which is $11.25 for a 60 second clip, and the Timeline 1.0 assembly surface at $0.10 per output minute, rounded up. Always read GET /v1/catalog for the live price of the route you call, since catalogs move.
The point is the comparison is not like for like. If your pipeline is mostly voice, compare Murf's quoted rate with the TTS price in GET /v1/catalog. If most of it is the video, the voice step is a small part of the bill.
Which should you choose?
Choose Murf for narration, dubbing and low-latency voice, where its quoted price and latency are its main selling points. Choose Sume when the unit of work is a video: script in, presenter video out, with captions and a job you can poll. Check the avatar video guide for the exact request fields.
Sources
Related posts
More in Comparisons
- Mux thumbnail time and fit_mode vs Sume video frames at and max_edge
Mux gets a thumbnail from a playback ID URL with time, width and fit_mode. Sume video-frames takes at[] times or fps and returns durable image files.
- Nano Banana 2: five characters, 14 objects vs Sume reference slots
Google says Nano Banana 2 keeps up to five characters and 14 objects consistent. Sume caps reference images per model; see what you can send and where it stops.
- Novita AI API alternative for video and image jobs: Sume
Novita offers model APIs, agent sandboxes and GPU deployment. Sume offers managed media jobs only. Where they overlap and where Novita does more.
- OpenRouter models fallback array and 3-entry limit vs Sume
OpenRouter's models array tries the next model on downtime, rate limits or moderation; fallbacks allows 3. Sume's allow_fallbacks has no effect.
Written by Sume