Rev AI forced alignment $0.003/min vs Sume caption script_text

Rev AI lists forced alignment at $0.003 a minute. Sume aligns your script to speech inside a $0.20 caption job. Which to use for subtitles from a script.

4 min readSume
All posts

If you already have the script and want timestamps, Rev AI lists forced alignment at $0.003 per minute, which is $0.18 for an hour of audio. Sume has no standalone alignment price: you send script_text on a caption job, Sume keeps the speech-to-text timings and aligns your text to them, and the job is $0.20 for a clip of up to 60 seconds. Rev gives you timing data cheaply; Sume gives you a captioned video.

What each product returns

The Rev AI page lists Forced Alignment at $0.003 per minute among its pay-as-you-go tiers. It is a timing service: you supply audio and text, and you get time stamps back. Sume's caption docs describe script_text as keeping speech-to-text word timings as the source of truth for time and aligning the burned-in text to your script, and the result is a rendered, captioned video.

Aligning a known script: Rev AI page vs Sume docs (read 2026-10-08)
QuestionRev AI forced alignmentSume caption job with script_text
Listed price$0.003 per minute$0.20 per job
30 seconds of audio$0.0015$0.20
60 seconds of audio$0.003$0.20
ReturnsAlignment dataCaptioned video
StylingNot part of the listingStyles such as slam, punch, tiktok-green plus design overrides

Why the difference exists

Per second, Sume is far more expensive, because it renders video. The right comparison is the whole pipeline: Rev alignment plus your own renderer plus hosting, against one request that returns the finished clip. If you have a renderer already, Rev's line wins. If not, the $0.20 includes the render.

Steps: caption with a known script

  • Host a clip of 60 seconds or less at a public HTTPS URL.
  • Send video_url, style and script_text to /v1/video-captions with an Idempotency-Key.
  • Handle the typed alignment errors the docs list if your script does not match the speech.
  • Download the captioned video from the job result.
curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: caption-script-001" \
  -d '{"video_url":"https://example.com/clean.mp4","style":"punch","script_text":"Say hello to the Sume developer platform."}'

What Sume does not do

Sume does not sell alignment as a separate data product and does not return a timing file from the caption job. You also cannot combine script_text with words, cues or segments in one request; the docs allow only one of the four.

When the script and the speech disagree

Alignment only works when the text matches the audio. The Sume docs list typed public job errors for a failed alignment, so build your retry around those: fix the script or drop script_text and let speech-to-text produce the words. Because failed jobs release the reserve, a mismatch costs you time rather than $0.20. With any alignment service, read the first result on a clean clip before you batch hundreds, and keep your script in the same language as the speech.

Language and style notes

The Sume caption docs separate style from language: language is only a speech-to-text hint and never selects the style or font. Korean text on a Latin style such as slam returns a 400 rather than rendering empty boxes. Rev AI's page lists forced alignment without style options, because it returns timing. Decide first whether you need a picture or only times.

Pick by deliverable

Need timestamps for your own editor: Rev AI at $0.003 a minute. Need a finished vertical clip with the script on screen: a Sume caption job at $0.20.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume