Edit a transcript with an AI instruction: Sume's STT flow
ElevenLabs STT accepts an edit instruction and returns edited_transcript. Sume's STT has no such field: it returns text and word timings to edit yourself.

Sume's speech-to-text cannot rewrite a transcript from an instruction. POST /v1/stt-1.0/transcribe returns text and word timings, and its request body is closed, so an extra instruction field is rejected. To edit, you transcribe first and apply the edit in your own step.
ElevenLabs facts are from its changelog; Sume facts from the API reference, both read 2026-09-30.
What did ElevenLabs add?
The September 28 changelog entry says batch and realtime speech-to-text accept a natural-language edit instruction of up to 2,000 characters, with the result returned in edited_transcript.
What fields does Sume's STT give me?
The request schema sets additionalProperties to false, so only the documented fields are accepted: audio_url, language_code, duration_seconds, segmentation, metadata and the usual mode, webhook_url and wait_timeout_seconds communication fields. Completed results expose text, language fields when available, and words[] as { word, start, end }.
With segmentation: { mode: "sentence" } you also get sentence segments derived from the returned word timings; the schema says it fails closed with a typed error if the provider returns no timed words.
| Step | ElevenLabs STT (changelog) | Sume STT (reference) |
|---|---|---|
| Edit instruction field | Up to 2,000 characters | None; body is closed |
| Edited text output | edited_transcript | Not returned |
| Original text | Not stated in the entry | text |
| Timings | Not stated in the entry | words[] always, optional sentences |
How do I edit a transcript on Sume?
Keep the words[] timings as the source of truth, make your text edit (by hand or with your own model), and map the changed words back to their start and end seconds. Timed words are what a cut list or subtitle file needs.
Two worked examples: edit subtitles, then burn them and a paper edit from a transcript.
Why keep the original text too?
An edit that rewrites wording no longer matches the audio. If the edited text drives captions, keep the original text and word timings so you can show what was actually said and re-time any change against it.
Sources
Related posts
- Auto-generate subtitles API: transcribe, fix the text, then burn
- Paper edit by API: build a rough cut from transcript lines
- Batch transcription API: transcribe many audio files
- Transcribe a video into sentence segments with Sume: options and cost
- Medical transcription API: what Sume STT does and doesn't
More in Developers
- 429 enforced_spend_limit_reached: why retrying never works
A 429 with no retry-after can be a monthly spend cap, not a rate limit. What Anthropic's page says, and how Sume separates 402, 429 rate_limited and queue_full.
- Face swap API: quality has no default, unlike avatar videos
Avatar videos default to plus when quality is omitted. Sume's Beta face swap has no default: quality is required and takes standard, plus or max.
- Firecrawl cache: maxAge 2 days vs Sume's fresh: true
Firecrawl scrape caches for 2 days by default (maxAge 172800000 ms). Sume's scrape has a boolean fresh, default false; fresh: true bypasses cached content.
- Firecrawl includePaths regex vs Sume's literal path prefixes
Firecrawl's includePaths are regex on the URL pathname. Sume's include_paths are literal prefixes starting with /, max 10 items. How to rewrite one.
Written by Sume