beehiiv audio block: make the MP3 from a video issue
beehiiv posts take mp3, wav, ogg, flac, aac and webm audio. Pull a 128 kbps MP3 out of your video issue with Sume's audio detach, $0.01 a job.

To add an audio version of a video newsletter issue to beehiiv, extract the track as an MP3 and drop it into an audio block. Sume's audio detach turns one workspace video into a new mp3 (128 kbps) or wav file for $0.01 a job, and beehiiv's support page lists mp3 and wav among the audio types a post accepts.
beehiiv facts here come from its support article on including audio files in a post, which shows an update date of August 27, 2026 and warns that it may not match the in-app experience (read 2026-10-02). Sume facts come from the audio detach docs.
What audio does a beehiiv post accept?
The article says a post accepts mp3, wav, ogg, flac, aac and webm. You insert the block by typing /audio in the editor and choosing Audio under Embeds, then upload a file or drag it in. The block shows a thumbnail, a title and a Play Online button.
In email, readers see that Play Online button, which sends them to the web version; on the publication website a full player appears with play and pause, skip and playback speed controls. You can restrict the block by website, email, anonymous viewers, free subscribers, paid tiers or referrals. The page states no file size limit, so check the editor before you upload a long file.
| Question | beehiiv audio block | Sume audio detach |
|---|---|---|
| Formats | mp3, wav, ogg, flac, aac, webm | wav (default, sample-exact) or mp3 at 128 kbps |
| Length | No limit stated on the page | Source up to 1800 s, output up to 900 s |
| In email | Play Online button to the web page | Not applicable, Sume only makes the file |
| Price | Not stated on the page | $0.01 per job (confirm live in GET /v1/catalog) |
How do I extract the MP3?
Import the video to Sume first with POST /v1/media-imports; audio detach takes only a media.sume.com URL from your workspace and an Idempotency-Key header. Set format to mp3. Add a range if you only want part of the track.
The call defaults to async, so you get a job back and read audio_url from the result:
curl -X POST https://api.sume.com/v1/audio-detach \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: issue-41-audio-mp3" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/issue-41.mp4",
"format": "mp3",
"range": { "start": 0, "end": 600 }
}'What goes wrong with a long issue?
A whole track past 900 seconds needs a range; the docs state that output is capped at 900 seconds, and a range longer than that is refused as audio_detach_range_empty. A 20 minute issue therefore becomes two jobs, each with its own range, and two audio blocks, or one block if you join them first with timeline audio concat at $0.01 flat per job.
If the video has no audio track, the job fails with detach_source_has_no_audio. Check probe.has_audio first with a video inspect call set to frames: false; probe is unbilled.
What should I check before publishing?
Sume does not post to beehiiv and has no beehiiv integration; you upload the file yourself. For the one-cent price and the range rule in more depth, see audio detach price and audio detach range over 900 seconds.
- Play the MP3 once. Sume re-encodes to 128 kbps; it does not normalize loudness.
- Keep the title in the beehiiv block short; the block shows a title next to the thumbnail.
- Confirm which audience tiers should hear it, because the block has its own access toggles.
- Remember that email readers get a button, not an inline player.
Should I pick MP3 or WAV?
Choose MP3 for the newsletter. beehiiv accepts both, but a WAV of ten minutes is far larger than a 128 kbps MP3, and email readers only get a Play Online button anyway, so quality beyond a spoken-word MP3 buys little. The Sume docs describe the default wav as sample-exact, which matters when audio will be joined again or used for lip-sync, not when it is the last stop.
Mono suits a talking issue. Set channels to mono to halve the data, and leave sample_rate out to inherit the source. If you plan to run speech-to-text on the same file later, the docs name 16000 Hz mono as the shape that speech-to-text wants, which is a separate wav job.
Sources
Related posts
More in Use cases
- Black Friday audio ad read: a TTS voice over a music bed
Produce a 30-second Black Friday ad read: write to character count, generate with Sume TTS, add a short music bed and mix. Costs and limits included.
- Build a 3-minute YouTube Short from clips with Sume Timeline
Stitch several clips into one Short of up to 180 seconds with Timeline 1.0: fades, a looped music bed with ducking, and the $0.30 render for three minutes.
- Bulk run says completed: re-queue only the failed holiday SKUs
A Sume bulk queue is completed once every item is terminal, not once all succeed. Read counts.failed, then re-queue only those under a new idempotency key.
- C2PA 2.3 live video support: does it reach AI clips made on Sume?
C2PA 2.3 adds live video and streaming support. Sume's video jobs return finished MP4 files, so the live feature changes little for generated clips.
Written by Sume