Newsletter issue as a 60-second avatar video and MP3
Turn a newsletter summary into a 60-second avatar video, then pull the MP3 for beehiiv's audio block with Sume's audio detach. What fits in a minute.
A newsletter can ship a spoken summary of each issue: generate a talking video of up to 60 seconds with Sume's avatar video, then pull the audio as an MP3 with audio detach for beehiiv's audio block. It is a summary, not a full read-aloud, because avatar scripts are limited to 4 to 60 seconds.
Sume's docs list no audio-only narration endpoint in the pages I read, so the voice here is the avatar video's. beehiiv's side is from its audio support article, updated August 27, 2026 per the page (read 2026-10-02).
What fits in 60 seconds?
Roughly the three headlines and one line of context each, plus a sign-off. Sume estimates the duration from the script and rejects anything outside 4 to 60 seconds, so write the summary first, submit it, and cut if it is refused. A full 2,000 word issue is far beyond the window and would have to be split into several jobs, which would sound like a list of fragments.
Pick the framing that suits the medium: the audio block is for listening, so avoid sentences that depend on a link or a chart. Say 'the table below' only if the reader can see it.
How do I produce the video and the MP3?
Submit the script to POST /v1/avatar-1.0/talking-video with an avatar_handle, then read the finished job's video_url, a media.sume.com artifact. Audio detach accepts that kind of URL directly, so no re-import is needed. Set format to mp3, which is 128 kbps; the default is wav.
Two calls, one for each step:
# 1. the video (async: poll /v1/jobs/:id/result for video_url)
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
-H "Authorization: Bearer $SUME_API_KEY" -H "Content-Type: application/json" \
-H "Idempotency-Key: issue-42-summary" \
-d '{"avatar_handle":"sume_clawra","aspect_ratio":"16:9","script":"This week: three things worth your time."}'
# 2. the MP3 from the finished video
curl -X POST https://api.sume.com/v1/audio-detach \
-H "Authorization: Bearer $SUME_API_KEY" -H "Content-Type: application/json" \
-H "Idempotency-Key: issue-42-mp3" \
-d '{"video_url":"https://media.sume.com/artifacts/artf_demo/issue-42.mp4","format":"mp3"}'How does the audio block treat it?
beehiiv accepts mp3, wav, ogg, flac, aac and webm, inserted with /audio under Embeds. Email readers see a thumbnail, title and Play Online button that opens the web version, while the site shows a full player. You can limit the block by audience, such as paid tiers only.
That means the MP3 reaches inbox readers one click away, not inline. If the summary is the hook, put its key point in the text above the block too.
| Step | Public rate in the docs |
|---|---|
| Avatar video, 4 to 60 seconds | By quality tier; see GET /v1/catalog |
| Audio detach to MP3 | $0.01 per job |
| beehiiv audio block | Not stated on the support page |
What should you disclose?
A spoken summary delivered by an avatar is synthetic. Tell subscribers in the issue text, in plain words near the audio block, and do not present the avatar as a real person on your staff. Rules on synthetic performers vary by place; the posts on California SB 1050 and audio detach pricing cover the two pieces this post touches. Track both jobs with Jobs and results.
What can go wrong?
The audio detach job fails with detach_source_has_no_audio if the video has no audio track, which should not happen for a talking video but is worth checking if you swap the source. A script outside 4 to 60 seconds is refused up front, and an audio file you export from the video keeps the avatar's voice, with no way to change that in the detach step.
Listen once before sending. Names, numbers and product terms are the usual trouble, so spell tricky words phonetically in the script if they sound wrong.
Keep the video too. The same job gives you a clip for the web version of the issue or social, so one script yields both the audio block and a short video.
Sources
Related posts
More in Use cases
- Nextdoor ad images: 1080x1080 or 2:1, 30 MB, 120-char headline
Nextdoor image ads use 1:1 at 1080 x 1080 or 2:1 at 1080 x 540, JPG or PNG up to 30 MB. Make the 1:1 with Sume and crop a 2:1 locally.
- Ofcom hash matching for AI deepfake intimate images: 30 September 2026
Ofcom said platforms should use hash matching to stop intimate-image abuse, including AI deepfakes, by 30 September 2026. What that means for AI video makers.
- Online course, 20 three-minute avatar lessons: cost by quality tier
Twenty 3-minute avatar lessons are 3,600 seconds. On Sume that is $662.40 Standard, $882.00 Plus or $1,980.00 Max, plus $0.95 once to create the avatar.
- OpenAI TTS requires AI-voice disclosure: burn it with Sume cues
OpenAI's TTS guide says to tell users the voice is AI-generated. A Sume caption job can burn that notice into video with authored cues, no ASR.
Written by Sume