Portuguese speech to text API: Sume STT language_code pt or pt-BR
Transcribe Portuguese audio with Sume STT using language_code pt or pt-BR, then check the reported language and word times. $0.01 per audio minute.

To transcribe Portuguese speech with Sume, submit the audio to POST /v1/stt-1.0/transcribe with language_code set to pt as a hint, wait for the job to complete, and read text and words[] from the result. Sume prices STT at $0.01 per audio minute, so a 90 second Portuguese clip reserves and bills about one and a half cents of audio time when you pass duration_seconds.
One honest caveat first. The Sume docs do not publish a list of supported languages for STT, so this page does not promise accuracy for Portuguese. It shows how to send the hint, how to check what came back, and how to judge a sample before you commit to a batch.
The request
Use pt as the hint for Portuguese. The field takes 2 to 16 characters, so pt-BR is a valid value for the schema, but Sume does not state that the region changes the result. Run the same clip with pt, with pt-BR, and with no hint, and keep whichever gives the text you would publish.
The job is asynchronous. Submit returns 202 with a request_id, which is the job id. Poll GET /v1/jobs/{id}/status until it is completed, failed or canceled, then read GET /v1/jobs/{id}/result. Always send an Idempotency-Key so a retried submit does not create a second paid job.
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: stt-pt-001" \
-d '{
"audio_url": "https://media.sume.com/artifacts/artf_demo/clip.wav",
"language_code": "pt",
"duration_seconds": 90,
"segmentation": "sentence"
}'
# then poll GET /v1/jobs/$REQUEST_ID/status and read /resultCheck the result before you trust it
Portuguese from Brazil and from Portugal differ in vocabulary, so a single sample is a poor test. Use a recording from the audience you serve. Read language_code in the result: if you sent pt-BR, see whether the job echoes pt-BR or pt, and store whatever it reports as the label for the transcript.
Compare the language_code in the result with the one you sent, and read language_probability. If they disagree, the clip may be mixed-language, very short, or mostly music. Run two or three of your own recordings, including a noisy one, and read the text yourself or have a speaker of the language read it.
| Field or fact | Where | Why it matters |
|---|---|---|
| language_code (request) | POST body, 2 to 16 characters | A hint such as pt-BR; omit it to let the service detect |
| language_code (result) | GET /v1/jobs/{id}/result | What the job reports it heard |
| language_probability | Same result | Low values mean review the text by hand |
| words[] | Same result | word, start, end in seconds for timing |
| Model page language count | microsoft.ai MAI-Transcribe-2 | 60 listed for that vendor model, not for Sume |
Captions from Portuguese speech
Portuguese uses the Latin alphabet with accents such as the cedilla and tilde, which is within the Latin font coverage the captions docs describe, but test one clip with accented words before a batch. If you want subtitles, send the STT text through video captions with script_text (up to 8000 characters) so Sume keeps the speech timings.
Cost and limits
One STT job takes at most 600 seconds of audio. For a longer recording, split it into slices, send one job per slice, and add each slice's start time to its word times when you merge them. Splitting a video's audio first with audio detach costs $0.01 per job and gives you a 16 kHz mono WAV.
If you leave duration_seconds out, Sume reserves one minute, which is wrong for a long file. Pass the real length rounded up. Microsoft's current model page lists 60 languages for its transcription model (read 2026-10-07), which is a vendor claim about that model and not a statement about Sume's STT.
The result also has segments[] when you ask for segmentation, and each segment carries start and end seconds. For subtitles you do not need to cut the text yourself: pass the transcript or your edited script to the captions endpoint and let it keep the speech timings.
Sources
Related posts
More in Developers
- Probe a finished video before upload: duration, size and aspect
Run video inspect with frames false to read a render's duration, size and frame rate before posting. Check it against the 3-minute Shorts limit.
- Python: cheapest Sume image model that lists your aspect ratio
A 25-line Python script reads Sume's image catalog, keeps models that list your aspect ratio, prices each from its endpoints record and prints the cheapest.
- Reconcile Sume jobs after a deploy or outage: poll what is open
After downtime, read status for every job your own table still shows as open, honor terminal and result_ready, and never resubmit. Python with sqlite.
- Redact faces and license plates: Pillow first, AI edit only to replace
For redaction use Pillow boxes you control; use an AI mask edit on openai/gpt-image-2.5 only to replace a plate or face, from $0.0094 per image on Sume.
Written by Sume