Telegram sendVoice formats: OGG/Opus, MP3 or M4A up to 50 MB
Telegram sendVoice accepts OGG/Opus, MP3 or M4A up to 50 MB. Pull an MP3 track from a generated video with Sume audio-detach and send it as a voice note.

To send a voice note from a bot, give sendVoice an OGG file encoded with Opus, an MP3 or an M4A, up to 50 MB. If your source is a generated video, the quickest route is Sume's audio-detach, which writes the soundtrack as a new MP3 or WAV artifact. Choose mp3, because it is on Telegram's accepted list, and send the resulting audio_url content through the bot.
Which audio formats does sendVoice take?
Telegram's sendVoice description names three: OGG encoded with OPUS, MP3 and M4A. It sets the upload ceiling at 50 MB for files sent this way. WAV is not on that list, which matters because it is one of the two output formats Sume's audio-detach can produce.
| Item | Telegram sendVoice | Sume audio-detach |
|---|---|---|
| Accepted or produced formats | OGG/Opus, MP3, M4A | wav or mp3 |
| Size ceiling | 50 MB | none stated; output capped by duration |
| Match for a voice note | mp3 | mp3 (128 kbps) |
| Price per job | not applicable | $0.01 |
How do I pull the MP3 with Sume?
Call POST /v1/audio-detach with a Sume-hosted video_url and format set to mp3. The call needs an Idempotency-Key header. It returns a job; poll GET /v1/jobs/:id/result and read audio_url, duration_seconds, format, channels and sample_rate. Optional range, channels and sample_rate fields let you take only the part you need, which also keeps the file small.
Videos from elsewhere must be imported first with POST /v1/media-imports, because Sume does not fetch arbitrary internet URLs.
import json, os, urllib.request
body = {
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"format": "mp3",
"channels": 1
}
req = urllib.request.Request(
"https://api.sume.com/v1/audio-detach",
data=json.dumps(body).encode(),
headers={
"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Content-Type": "application/json",
"Idempotency-Key": "voice-note-001",
},
method="POST",
)
print(urllib.request.urlopen(req).read().decode())What about the 50 MB limit and length?
The size ceiling is rarely the issue for speech: Sume encodes MP3 at 128 kbps, which is about 16 KB per second, so a 50 MB cap sits far beyond a chat clip. What can bite is duration, because Sume's audio-detach caps the output length, so trim a long video first with video trim when you only want a section.
What should I verify before shipping?
Confirm the file plays in a test chat on the clients you care about. Telegram's page lists MP3 as accepted, so the method should take it, but only a real send proves how each client shows it. The audio-detach guide lists every field and refusal code.
Sources
Related posts
More in Integrations
- TikTok max_video_post_duration_sec: trim before Direct Post
TikTok's creator_info endpoint returns max_video_post_duration_sec, a per-creator limit. Read it first, then cut the clip to fit with Sume video trim.
- TikTok privacy_level_options: public vs private account values
TikTok creator_info returns different privacy_level_options for public and private accounts, and Direct Post privacy_level must match one. Values for each.
- TikTok publish webhooks vs Sume run webhooks: wire both safely
TikTok sends webhooks for failed, complete, inbox, public and removed posts. Sume signs its own run webhook separately. Keep two receivers and verify Sume's.
- Val Town free 1-minute timeout: a Sume webhook receiver val
Val Town's free plan stops a val at 1 minute and runs crons every 15 minutes at best. Submit Sume jobs async, then take the result by webhook or a slow cron.
Written by Sume