How do I narrate a family history video with AI voice and photos?

Narrate a family history from old photos: 1,800 characters of script is $0.09 on Sume TTS, plus a $0.125 music bed and a $0.30 Timeline render.

4 min readSume
All posts

To narrate a family history video, scan the photos, write a script of about 1,800 characters, generate it as a TTS job, add a quiet music bed, and render it over the stills with Timeline 1.0. On Sume the voice is $0.09, the bed $0.125 and the render $0.10 per started minute, so a 150-second video is $0.515 with three started render minutes.

The photographs carry the emotion. The voice's job is to be warm, unhurried and accurate, and to leave silence where a face needs time.

Write what the photos cannot say

Do not describe the picture. Say who, when and why it matters: the street, the job, the sentence someone always said. Keep one idea per photo and one photo per 8 to 15 seconds. Cartesia's pricing page puts a minute of speech at 750 to 800 credits, so 1,800 characters is about two and a half minutes. Cut to 1,500 for a two-minute film, or keep 1,800 and leave the photos on screen for the pauses.

Names are the hard part. A surname read wrongly in a family film is painful. Add a pronunciation dictionary for each name, listen to every name alone before generating the whole script, and spell uncertain names as they sound in the dictionary entry. Sume TTS accepts a pronunciation_dict_id on the request.

Choose a voice and pace

A slower voice suits memory. Set generation_config.speed to 0.85 to 0.95, ask for a gentle emotion in the free-text field, which is limited to 64 characters, and avoid a bright announcer tone. If the family speaks another language at home, send the matching language code and a voice made for it; a mismatch returns a 409 tts_voice_language_mismatch before any charge.

A cloned voice of a relative is a different question, with consent and licensing at its centre; this post covers a neutral narrator only.

Prepare the photos

Scan at a resolution well above the video size and crop to the frame you will render, so nothing is stretched. Put the photos in the order of the script and name them 01, 02, 03 so the slots match. If a photo is damaged, repair it before it goes into the render, since the render uses the file as given.

Group portraits need longer on screen than landscapes. Give each a duration that matches the sentence it illustrates, and add a second or two of picture with no speech at the start and the end of the film.

Pull it together on the timeline

Import the scans, generate the voice as wav, and send both to a Timeline 1.0 render. Each photo is a still in a video[] slot; stills are allowed, so no clip generation is needed. The audio.duration_seconds must match the voice file and the slot durations should add up to it. Add a soft, instrumental bed from Music Router as the soundtrack, with a duck_db of about 10 so it sits under the speech and a fade_out_seconds of 4 or 5 for the ending.

curl -X POST https://api.sume.com/v1/timeline-1.0/render \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: family-history-001" \
  -d '{
    "audio": { "url": "'"$VOICE_URL"'", "duration_seconds": 150 },
    "video": [
      { "source_url": "'"$PHOTO_1"'", "start": 0, "duration": 50 },
      { "source_url": "'"$PHOTO_2"'", "start": 50, "duration": 50 },
      { "source_url": "'"$PHOTO_3"'", "start": 100, "duration": 50 }
    ],
    "soundtrack": { "url": "'"$BED_URL"'", "gain_db": -10, "duck_db": 10, "fade_out_seconds": 5 },
    "output": { "width": 1920, "height": 1080 }
  }'

Cost and keeping the files

The render is metered per started output minute, so 150 seconds is three minutes at $0.10 each. If a name is corrected later, retake only that sentence and splice it into the narration with a concat, then render again; the voice is a cent or two and the render is the main cost.

Keep the scans, the script, the pronunciation dictionary and the finished audio together in one folder. Relatives will ask for a second version.

A 150-second family history on Sume, catalog rates as of 2026-10-07; vendor voice list rates read 2026-10-07 for 1,800 characters.
ItemRateCost
Sume TTS 1.0, 1,800 characters$0.0475 per 1,000 characters$0.09
Music Router bed$0.125 per generation$0.125
Timeline 1.0 render, 3 started minutes$0.10 per started output minute$0.30
For comparison: MAI-Voice-2.1 Standard voice only$22 per 1M characters$0.040
For comparison: ElevenLabs Flash/Turbo voice only$0.04 per 1,000 characters$0.072

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume