Maestra translates 30 minutes for $12: what Sume offers instead
Maestra's $12 for 60 credits buys 30 minutes of translation. Sume has no translate endpoint, so a dub is STT, your own translation, TTS and a timeline render.

On Maestra's pricing page 60 credits cost $12 and cover either 60 minutes of subtitles or 30 minutes of translation. Sume does not translate text, so the equivalent job is a pipeline: detach the audio, transcribe it, translate with a tool of your choice, synthesize speech and render the result.
The Sume pipeline, step by step
Each step is one documented endpoint with its own price, so you can see where the money goes.
| Step | Endpoint | Price |
|---|---|---|
| Detach audio | POST /v1/audio-detach | $0.01 per job |
| Transcribe | POST /v1/stt-1.0/transcribe | about $0.01 per minute |
| Translate | Your own tool | Not a Sume product |
| Speech | POST /v1/tts-1.0/generate | Per character; check the catalog |
| Render | POST /v1/timeline-1.0/render | $0.10 per output minute |
What the numbers say
Excluding translation and speech, a 10-minute video costs about $1.11 in the Sume steps: $0.01 to detach, about $0.10 to transcribe, and $1.00 to render. Add the translation tool and the per-character TTS price from GET /v1/catalog for the full figure.
Maestra's translation is $0.40 per minute from the page's own arithmetic, bundled with the translation itself. Sume's total depends on the translator you pick, which is the trade for controlling the wording.
When to build it yourself
Build it when you need translation review, a glossary, or a particular voice. STT takes duration_seconds up to 600 per job, so a longer file is transcribed in chunks.
Sources
Related posts
More in Comparisons
- MAI-Transcribe-2 'one hour in 10 seconds': a short caption clip
Microsoft lists 1hr audio to 10 sec for MAI-Transcribe-2. What that ratio omits for a 20-second clip, and the Sume STT, detach and caption limits that apply.
- MAI-Transcribe-2-Streaming partials vs Sume STT word timings
MAI-Transcribe-2-Streaming sends partials after about 100 ms in 60 languages. Sume STT is a job that always returns word timings. Which fits what.
- Make an AI avatar say exact words: Tavus echo mode vs Sume
Tavus echo mode sends text or audio straight to the avatar for playback, skipping perception and speech recognition. Sume's avatar video renders your script.
- Micro-drama lead: Avatar 1.0 or Seedance 2.5? Pick by dialogue
Pick Sume Avatar 1.0 for talking to camera and Seedance 2.5 for a lead who moves through places. Cost per minute: $14.70 against $34.67.
Written by Sume