Press release video: an announcement read by an AI avatar

A press release video reads the headline, key facts, and a quote in about a minute, captioned from the release text. How to make one with an AI avatar.

5 min readSume
All posts

A press release video is a short video version of a company announcement: a presenter reads the headline, the key facts, and one quote in about a minute, and captions in the release's own wording carry it for viewers watching without sound. It sits on the newsroom page next to the written release and in the social posts that link to it.

With an AI avatar, the approved release text becomes the script, so the video can go out with the release. On Sume, one avatar reads the script as a talking video, a caption job burns in the same wording, and a longer release is split and joined. Facts come from Generate avatar video and Video captions, read on 2026-09-28, plus Sume's current code where noted.

What should a press release video script include?

Cut the release down to what a viewer needs in one sitting; the full text stays on the page.

  • The headline, in the first sentence.
  • Who, what, when, and where, in two or three sentences.
  • One quote, with the speaker's name and title said aloud.
  • Where to read the full release or contact the press team.

How do I turn a press release into an avatar video?

Create one presenter for company news and keep its handle, then send the script to the talking-video route. Use 16:9 for the newsroom page and 9:16 for vertical social posts; 9:16 is the default, and each shape is its own render. A scene prompt sets the room. Completed results can include public media.sume.com video files, so treat the URL as public and keep it out of anything shared before the release goes out.

curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: press-second-factory-16x9" \
  -d '{
    "avatar_handle": "newsroom_host",
    "script": "Today Acme opened its second factory. Here is what that means for customers.",
    "aspect_ratio": "16:9",
    "scene": { "type": "prompt", "prompt": "Neutral newsroom set" }
  }'

How do I add captions from the release text?

Send each finished clip to POST /v1/video-captions with its video_url and the script as script_text, so Sume aligns the burned-in wording to your script, not to what speech recognition heard. In current code the caption job refuses a source longer than 60 seconds or without an audio stream, so caption each clip before any join. Burning captions onto a video covers the request, the styles and the alignment errors.

What if the release needs more than a minute?

One talking video covers an estimated 4-60 seconds, so split a longer script at paragraph breaks and render each part with the same handle and shape. Timeline 1.0 joins the parts into one MP4 of up to 1,800 seconds. In current code a Timeline render's sound comes only from its audio spine and an optional soundtrack, so the speech has to be carried on the spine; making an avatar video longer than 60 seconds shows the join. For a bulletin of several stories, see the AI news anchor post.

What are the limits, and what does it cost?

  • In current code the avatar route speaks English only; what languages an AI avatar can speak covers releases in other languages.
  • resolution is currently 720p, with one presenter per video.
From Generate avatar video, Video captions, Timeline 1.0 and API pricing, read 2026-09-28. Prices are before a 5.5% agent fee by default.
PieceLimitPrice
Release clipEstimated 4-60 seconds$0.184/s standard, $0.245/s plus, $0.55/s max (no product image)
Caption jobOne clip up to 60 seconds$0.20 per job
Join parts (optional)Up to 1,800 seconds$0.10 per output minute

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume