How do I make AI announcer intros for conference speakers?
Ten 25-second speaker intros with one shared sting cost about 43 cents on Sume: 2 cents of TTS and a 1-cent join each, plus a $0.125 track made once.

Make speaker intros for a conference by writing a 300-character announcement per speaker, running one TTS job each, and joining every file to one shared sting with a Timeline audio concat. Each intro is about 25 seconds, costs $0.02 of TTS plus $0.01 for the join, and the sting is a single $0.125 Music Router track made once. For 10 speakers that comes to $0.425 on Sume at the catalog rates of 2026-10-07.
This is for pre-recorded playout from a laptop at the venue. It is not live: Sume's TTS returns finished files from asynchronous jobs, so everything is rendered before the doors open, and a speaker who swaps in at the last minute costs one more job and one more concat, not a new session.
Write the intro to be heard once
A stage intro has three jobs: say who is coming, say why the audience should care in one line, and stop. Cartesia's pricing page gives one minute of Sonic speech as 750 to 800 credits at one credit per character, the same metering Sume TTS 1.0 uses, so 25 seconds is about 310 characters. Use that as a ceiling.
At $0.0475 per 1,000 characters, 310 characters is $0.0147, rounded up to $0.02. Put the name late in the first sentence, spell hard names the way they are said, and keep titles short. 'Please welcome Marta Nowak, who leads the grid team at ExampleCorp' is easier to say well than a full job title and the company's legal name.
Names are the risk, so test them first
The failure that embarrasses a speaker is a mispronounced name. TTS 1.0 accepts an optional pronunciation_dict_id, which lets you attach a pronunciation dictionary to a request, and the language field sets the language the voice reads in. For a speaker whose name is from another language, the cheapest check is a one-sentence job with just the name, which costs a cent, and a listen by the speaker or a colleague who knows it.
Send each speaker the audio and ask for approval. A speaker who hears the intro two days ahead rarely objects on the day.
Build the sting once, join it to each intro
Music Router takes a prompt of up to 5,000 characters and rejects a duration field, so ask for the length in the prompt: 'A 6-second confident stage sting, 120 BPM, bright brass and a tight kick, ending on a clean stop. Instrumental, no vocals.' Every router model charges the same $0.125 per generation. The join then puts the sting first and the intro after it, or the other way around.
The concat takes up to 20 parts from your workspace's hosted audio, no re-synthesis, no silence at the seams. Use wav for the voice if you will join it, since mp3 adds encoder padding at each edge.
curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: stage-intro-nowak-001" \
-d '{
"operation": "concat",
"parts": [
{ "url": "https://media.sume.com/artifacts/artf_demo/stage-sting.wav" },
{ "url": "https://media.sume.com/artifacts/artf_demo/intro-nowak.wav" }
],
"output": { "format": "wav" }
}'Voice over the sting, or after it
The sequence above plays the sting and then the voice. If you want the voice speaking over the music, the documented route is a Timeline render: the intro as the audio spine, the sting as the soundtrack with gain_db, loop and duck_db, and one still image as the only video slot. The result is an MP4 at $0.10 per started output minute, so $0.10 per intro instead of $0.01. For a venue laptop that plays audio from MP4 files, that is fine; if you need a plain audio file, sequence it.
- Generate every intro a week ahead and keep the job ids with the speaker list.
- Keep a spare generic intro for no-shows and walk-ons.
- Test levels in the actual room; the file is the same, but the PA is not.
| Approach | TTS (10 x 310 chars) | Sting | Join or render | Total |
|---|---|---|---|---|
| Sting then voice, audio concat | $0.20 | $0.125 | $0.10 | $0.425 |
| Voice over sting, MP4 render | $0.20 | $0.125 | $1.00 | $1.325 |
One host voice, more than one language
Keep the same voice id for every speaker so the stage sounds like one host. If the event is bilingual, run each intro twice with the right language and a voice that carries it; a mismatch returns a 409 before any charge, so a wrong pairing costs you nothing but a retry. Budget the second language as a second set of jobs at the same price per intro.
Sources
Related posts
More in Use cases
- How do I make a safety briefing audio in three languages with TTS?
One 1,800-character briefing in English, Spanish and Polish is three TTS jobs: 9 cents each, 27 cents in all on Sume. Set language on every job and review it.
- How do I narrate a family history video with AI voice and photos?
Narrate a family history from old photos: 1,800 characters of script is $0.09 on Sume TTS, plus a $0.125 music bed and a $0.30 Timeline render.
- What does a first-time home buyer video series cost with an avatar?
Six 30-second avatar tips cost $33.12 at standard, $44.10 at plus or $99.00 at max on Sume, plus $0.95 once for the avatar. Rates and script limits.
- Four seasons from one venue photo: four Ideogram 4.5 edits for $0.30
Make spring, summer, autumn and winter versions of one venue photo with four ideogram/ideogram-v4.5 edits at $0.075 each on Sume, $7.50 for 25 venues.
Written by Sume