How do I make a self-guided walking tour audio with TTS?
A six-stop walking tour at 700 characters a stop is six TTS jobs of 4 cents each, 24 cents on Sume, about 0.9 MB per stop as default mp3.

For a self-guided walking tour, write one script per stop, generate one TTS file per stop, and let your app play the right file when the walker reaches that stop. Six stops at about 700 characters each is six jobs at $0.04 each, $0.24 on Sume, and each default mp3 is roughly 0.9 MB. The job is asynchronous, so generate everything before the tour goes live and ship the files with the app or from your own storage.
You can also send all six scripts as one 4,200-character request ($0.20) and split the result with sentence timings and a Timeline audio split, but one file per stop is simpler to update: a corrected stop is one re-run, not a re-split.
Size each stop for a walker
Cartesia's pricing page says one minute of Sonic speech takes 750 to 800 credits at one credit per character, and Sume TTS 1.0 uses one credit per character. At 750 characters a minute, a 700-character stop plays for about 56 seconds, which is long enough to say something worth stopping for and short enough that nobody stands in the rain. Keep the opening sentence about what to look at, not about history, so the audio matches what the walker sees.
Slow the narrator slightly with generation_config.speed at 0.95 or 0.9; the range is 0.6 to 1.5 and the setting changes the length of the audio, not the price, because the bill is by character.
File size and format
The default output is mp3 at 44.1 kHz and 128 kbps. At 128,000 bits a second, that is 16,000 bytes a second, so a 56-second stop is about 0.9 MB and six stops are about 5.4 MB. If the app has to be small, output_format accepts lower rates: a 64,000 bits-per-second mp3 halves the size at some cost in fidelity, and speech at 22.05 kHz is plenty for a phone speaker. The documented sample rates are 8,000, 16,000, 22,050, 24,000, 44,100 and 48,000 Hz.
| mp3 bit rate | Bytes per second | One stop (56 s) | Six stops | TTS price for six stops |
|---|---|---|---|---|
| 128 kbps (default) | 16,000 | 0.90 MB | 5.4 MB | $0.24 |
| 96 kbps | 12,000 | 0.67 MB | 4.0 MB | $0.24 |
| 64 kbps | 8,000 | 0.45 MB | 2.7 MB | $0.24 |
Generate the stops in one loop
Each stop has its own idempotency key built from the stop number, so a rerun of the script returns the existing job for stops you have already paid for. The script below submits six jobs and prints the job ids; collect the audio artifact for each from the job result once it is completed, or let a webhook tell you.
const stops = [
"Stop one. The old customs house. Look up at the clock above the door.",
"Stop two. The fish market. Notice the iron roof on your left.",
];
for (const [i, text] of stops.entries()) {
const res = await fetch("https://api.sume.com/v1/tts-1.0/generate", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.SUME_API_KEY}`,
"Content-Type": "application/json",
"Idempotency-Key": `walk-tour-stop-${i + 1}`,
},
body: JSON.stringify({
transcript: text,
voice: { id: process.env.VOICE_ID },
language: "en",
generation_config: { speed: 0.95 },
output_format: { container: "mp3", sample_rate: 24000, bit_rate: 64000 },
mode: "async",
}),
});
const out = await res.json();
console.log(i + 1, (out.data ?? out).request_id);
}What to add, and what to leave out
An ambient bed is optional and costs $0.125 for one Music Router track, which you can loop in your own player; Sume's audio concat joins files but does not layer them, and layering is only a Timeline render that produces an MP4. Directions between stops can be a seventh short job. Do not rely on the audio to be the only warning about a hazard, and do not tell people to cross a road in a recording, because conditions change and a file does not.
If the tour has a second language, run the same six jobs again with that language's script, the language field set and a voice in that language. A voice and language mismatch returns a 409 before any charge, so the first sentence of the second run tells you whether the voice fits.
Budget for revisions as well as the first pass. A tour with three languages is eighteen jobs, $0.72 at the same character counts, and each corrected stop adds $0.04 for one language. Because every stop is a separate file, a council that renames a street changes one job, not the whole tour. Keep each stop's script, language, voice id and job id in one table so a reviewer can find the file that needs a rerun in seconds.
Sources
Related posts
More in Use cases
- How do I add a spoken tagline and sound sting to the end of an ad?
Make a 3-second spoken end tag: one 1-cent TTS job, a $0.125 music sting and a $0.01 concat to attach it to any ad read. About 14.5 cents on Sume.
- Swap the presenter per market: one ad, two Recast jobs on Sume
Recast keeps the motion, camera, cuts and sound of one source ad and replaces the people. How to make one version per market, the limits and the price.
- 10 static ad variants in one Sume image request: n and its cap
The Sume Image API takes n from 1 to 10, but each model has its own cap. Read it from the catalog; use 4:5 on Nano Banana or 1088 x 1360 on GPT for feed ads.
- How do I put three products in one AI video from reference photos?
Send three product photos as input_references to Gemini Omni Flash 1.1: one 8-second 720p clip is $1.00 on Sume, against $1.89 for three single clips.
Written by Sume