Museum audio guide API: one TTS job per exhibit, about $0.08

Write one 1,500-character script per exhibit, run each through Sume TTS with a language and pronunciation dictionary, and keep sentence timestamps.

4 min readSume
All posts

Treat each exhibit as one TTS job. A 1,500-character script costs about $0.071 on Sume's rate of $0.0475 per 1,000 characters, which rounds up to $0.08 for the job, so a 40-exhibit guide runs about $3.20 per language.

Why one job per exhibit

A TTS request takes 1 to 20,000 characters and fails if the audio would run past 1,200 seconds. Per-exhibit jobs stay far under both limits, let visitors jump to any stop, and mean that rewriting one label regenerates one file. Set language on each job, and confirm_language_mismatch only when you deliberately read a text in a language the voice does not match.

Audio guide cost by exhibit count (read 2026-10-03)
ExhibitsCharacters eachPer exhibitTotal per language
101,500$0.08$0.80
401,500$0.08$3.20
1001,500$0.08$8.00

Names and captions

Artist and place names are where TTS fails first. A pronunciation_dict_id on the request applies a dictionary you created, so the same spelling is read the same way in every exhibit. For on-screen text, ask for timestamps.words with segmentation.mode: "sentence". Segmentation requires word timestamps to be on, or the request is refused.

Audio format for kiosks

The default output is mp3 at 44,100 Hz and 128 kbps. Choose wav when a player needs uncompressed files, or when you want per-sentence audio from emit_audio, which needs a wav or raw container. Name the files after your own exhibit ids and store the job metadata in the metadata field. The sentence-cut post shows the slicing.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume