How do I make a trade-show booth audio loop with AI voice and music?

Make a 60-second booth loop from one TTS job, one music bed and a Timeline render with a looping soundtrack: about 25 cents on Sume, then play it on repeat.

4 min readSume
All posts

To make a booth audio loop with an AI voice and music, write a 60-second script, generate it as one TTS job, generate a music bed, and combine them in a Timeline 1.0 render that uses a still or short clip as the picture. At 700 characters the voice is $0.04, the bed is $0.125 and the render is $0.10 for one output minute, $0.265 in all. The finished MP4 plays on repeat on any screen or laptop at the stand.

Booths are noisy and people pass at three-second intervals. The loop's job is to say who you are in the first sentence, give one reason to stop, and end on a line that restarts cleanly.

Write the script as a loop

Because the file repeats, its last line should lead into its first. 'Come ask us how' followed by a restart on 'Sume turns text into video' reads as one thought. Avoid a closing sign-off that sounds final; a listener who joins at second 40 should not hear 'thanks for listening'.

Keep sentences short and repeat the name of the product twice in 60 seconds. At roughly 750 characters a minute of speech, per Cartesia's pricing page, 700 characters gives you a few seconds of breathing room, which is useful in a loud hall where listeners need gaps.

Generate the voice and bed

Use a clear, mid-paced voice, set generation_config.speed to about 0.95, and ask for wav output so nothing is re-encoded before the render. For the bed, Music Router takes one prompt of up to 5,000 characters and charges $0.125 per generation whichever model it routes to. Put the length in the prompt, because duration is rejected, and ask for an instrumental with a steady pulse and no long builds, since a build is the part that sounds odd when repeated.

curl -X POST https://api.sume.com/v1/music-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: booth-bed-001" \
  -d '{
    "model": "sume/music-auto",
    "prompt": "A 60-second upbeat corporate-tech instrumental, 110 BPM, steady pulse, light synth arpeggios, no long builds, no vocals, ending on a resolved chord."
  }'

Render it with a looping soundtrack

The Timeline 1.0 soundtrack object takes a loop flag, gain_db, a duck_db of 0 to 20 and fade_out_seconds up to 10. If the bed is shorter than the voice, loop it; set a duck of 10 to 14 dB so the words sit above the music even when the hall is loud. Use a single still of your logo or a key product shot as the video slot, since stills are allowed, and set audio.duration_seconds to the voice file's length.

curl -X POST https://api.sume.com/v1/timeline-1.0/render \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: booth-render-001" \
  -d '{
    "audio": { "url": "'"$VOICE_URL"'", "duration_seconds": 58 },
    "video": [{ "source_url": "'"$LOGO_STILL_URL"'", "start": 0, "duration": 58 }],
    "soundtrack": { "url": "'"$BED_URL"'", "gain_db": -10, "duck_db": 12, "loop": true },
    "output": { "width": 1920, "height": 1080 }
  }'

Test it where it will play

A script that works in headphones can vanish on a stand. Listen to the first render on the speaker you will use, from the distance a visitor stands, with someone talking nearby. If a word is lost, the fix is usually in the script: shorter words, a slower speed value, a deeper duck on the bed. Rewrite, regenerate the voice only, and render again.

Make two versions if the show runs several days: one with the current offer and one with a different opener. The bed is reusable, so the second version costs a TTS job and a render, 14 cents at this length: 4 cents of speech and the $0.10 render. Keep both files and note which ran on which day, so you can tell which opener stopped more people.

Cost and playback

A render is metered per started output minute, so a 58-second loop is one minute. Revisions are cheap enough to do the night before the show: a new voice take is another TTS job, and a new render is another $0.10.

Playback is the dull part. Export once, test the file on the exact screen and speaker you will use, and set it to repeat in the player. Mind the volume rules of the venue, and keep a text version of the script on the screen in case the sound is switched off.

Cost of a 60-second booth loop on Sume, catalog rates as of 2026-10-07; vendor rate read 2026-10-07 from the Cartesia pricing page for the 750 characters per minute assumption.
ItemUnitPriceThis loop
TTS 1.0, 700 charactersper 1,000 characters$0.0475$0.04
Music Router bedper generation$0.125$0.125
Timeline 1.0 renderper started output minute$0.10$0.10
Total$0.265

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume