Text to speech from a PDF with an API: extract, chunk, narrate
Sume's TTS takes text, not PDF files. Extract the text, split it under 20,000 characters, narrate each chunk, then join the audio. Costs and limits included.

To turn a PDF into speech with Sume, extract the text yourself, split it into chunks of 20,000 characters or fewer, send each chunk to the text-to-speech route, and join the finished audio with timeline audio. Sume's TTS request accepts a transcript string, not a PDF, a URL or an upload, so the extraction step is yours.
Everything about Sume below comes from the Sume API reference, the timeline audio guide and the jobs and results guide.
Why does Sume not read a PDF directly?
The request schema has two text inputs, transcript (1 to 20,000 characters) and transcript_source (a reference to an accepted script), and exactly one is allowed. There is no file field. That keeps the price predictable, because billing is per character of the text you send, with spaces and punctuation counted.
It also means you decide what gets read. A PDF often contains page numbers, running headers, footnotes, hyphenated line breaks and tables. A voice will read all of them if you leave them in, so clean the text before it reaches the API.
How do you prepare the text?
Do these steps with any extraction tool you trust, before calling Sume:
- Remove running headers, footers and page numbers.
- Join words hyphenated across lines and collapse hard line breaks inside paragraphs.
- Decide what to do with tables, figures and footnotes: summarize them in a sentence or leave them out.
- Spell out symbols and abbreviations the voice may misread, and write numbers the way you want them spoken.
- Split at paragraph or heading boundaries so each chunk is under 20,000 characters and each result stays under the 1,200-second audio cap.
What does the narration call look like?
Send one request per chunk. Use the same model, the same voice and the same language on every chunk so the document sounds like one reading, and give each request its own idempotency key such as the chunk number. For non-English text set language; a known mismatch between the voice and the language returns 409 tts_voice_language_mismatch before any charge.
curl -X POST https://api.sume.com/v1/tts-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: narration-001" \
-d '{
"model": "sonic-3.6",
"transcript": "Welcome back. Today we compare two ways to get a voiceover.",
"voice": { "id": "'"$VOICE_ID"'" },
"language": "en"
}'
Poll each job, then take audio_url from the result:
curl https://api.sume.com/v1/jobs/$JOB_ID/status \
-H "Authorization: Bearer $SUME_API_KEY"
curl https://api.sume.com/v1/jobs/$JOB_ID/result \
-H "Authorization: Bearer $SUME_API_KEY"
How do you join the chunks and what does it cost?
The default output is mp3 at 44.1 kHz and 128 kbps. If you plan to join and edit, ask for a wav container, then import any files that are not already Sume-hosted. Timeline audio concat takes 1 to 20 Sume-hosted parts, joins them without gaps at the seams and costs $0.01 per job.
A 40-page report with about 100,000 characters comes to 100 x $0.0475 = $4.75 of speech at the catalog price, in five or more chunk jobs, plus a concat. That is arithmetic on the catalog price, so check the live catalog before budgeting. Each retake of a chunk is billed as a new job, so listen to chunk one first, then run the rest.
| Item | Value |
|---|---|
| Accepted input | transcript text, 1 to 20,000 characters per request |
| PDF or file upload | Not accepted by the TTS route |
| Audio per job | Up to 1,200 seconds |
| Price | $0.0475 per 1,000 characters at the catalog price |
| Parts per concat job | 1 to 20, $0.01 flat per job |
What does this not do?
Sume does not extract text, run OCR on scanned pages, summarize a document or keep page-level timing. If the PDF is a scan, you need OCR first, and OCR errors will be read aloud faithfully. For word timings, request timestamps.words on each chunk and offset them by the start time of each part in the concat result. Keep the original document and the cleaned text side by side so a listener's correction can be traced to a source line.
Sources
Related posts
More in Use cases
- Thanksgiving recipe video: step cards with caption cues
Make a 40-second Thanksgiving recipe video from five phone shots: join them in Timeline 1.0, burn numbered step cards with caption cues, add a music bed.
- TikTok in-feed ad caption rules: no hashtags, @ or clickable links
TikTok non-Spark in-feed ad captions are white in a fixed font, cannot hold links, @ symbols or hashtags. Where a burned-in caption fits instead.
- TikTok ad disclaimers: 90 characters, 3 links, and the AI type
TikTok Ads Manager has three disclaimer types: 90-character text, up to 3 clickable links, and an AI-generated label. What each does and what Sume does not set.
- TikTok video package opening and end scenes: 1-10 seconds, 720x1280
TikTok's video package page lists opening and end scenes of 1 to 10 seconds in MP4, MOV, FLV, MKV or WebM, and Lottie templates under 16 s. Trim with Sume.
Written by Sume