Talking photo audio too large? The 10 MB Sume Fabric limit
Sume's veed/fabric-1.0 needs Sume-hosted audio under 10 MiB and up to 300 seconds. Size arithmetic for WAV and MP3, and how to split or compress.

Sume's talking-photo route, POST /v1/veed/fabric-1.0, accepts audio that is hosted on Sume (the media host) and at most 10 MiB (10,485,760 bytes), for up to 300 seconds. A file over the byte cap is rejected with 400 audio_too_large. If yours is, split it into shorter parts or re-export it smaller before you submit. A 16-bit stereo WAV at 44.1 kHz is about 176 KB per second, so it reaches 10 MB in under a minute; that is the usual cause.
The arithmetic below uses 10,000,000 bytes, slightly under the real cap, so it errs on the safe side. It is a rule of thumb, not a Sume-published table.
How long can an audio file be?
Uncompressed audio is bytes per second times seconds.
| Format | Bytes per second | Seconds in 10 MB |
|---|---|---|
| WAV 16-bit stereo, 44.1 kHz | 176,400 | About 56 |
| WAV 16-bit mono, 44.1 kHz | 88,200 | About 113 |
| WAV 16-bit mono, 16 kHz | 32,000 | About 312, so the 300 s cap applies first |
| MP3 at 128 kbps | 16,000 | About 625, so the 300 s cap applies first |
Three ways to get under the limit
Every route needs the audio on media.sume.com. Import an outside file with POST /v1/media-imports, or use the URL from a previous Sume job.
- Mono and a lower sample rate. Speech does not need stereo; a mono 16 kHz WAV fits the full 300 seconds.
- Split the audio. Timeline audio can slice one Sume-hosted file into up to 20 ranges for $0.01 a job, and each range becomes its own file.
- Use MP3 for the upload. The timeline audio docs warn that MP3 re-adds encoder padding at every edge, so prefer WAV when the file will be joined again or drives lip-sync, and use MP3 only when size is the problem.
Splitting example
Cut a long take into two parts, then run each through Fabric with its own duration.
curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: split-narration-001" \
-d '{
"operation": "split",
"url": "https://media.sume.com/artifacts/artf_demo/narration.wav",
"ranges": [{"start": 0, "end": 45}, {"start": 45}]
}'Checking size before you send
Check the byte size of the file before upload rather than waiting for a rejection. Most shells give it in one command, and a script can refuse anything near 9 MB so you have a margin for headers and for rounding. If the audio came from a text-to-speech step, export it as mono at the lowest sample rate your listeners will not notice, which is usually plenty for speech and keeps files small without splitting.
Limits
Fabric is billed per second of audio, rounded up, at $0.1875 at 720p and $0.10 at 480p. Splitting does not change the total except for rounding: each part is rounded up to a whole second. Two parts of 45.2 seconds bill as 46 seconds each, one second more than a single 90.4-second job would bill as 91. Send a measured duration_seconds that matches the audio.
Sources
Related posts
More in Media tools
- Which AI video model makes 1:1 square clips on Sume?
Kling 3.0, Wan 3.0, MiniMax H3, H3 Max and Grok Imagine list 1:1 in Sume's Videos panel; Auto does not. What to pin for square feed video.
- Captions with sound cues for deaf viewers: W3C checklist on Sume
W3C says captions carry speech and non-speech sound. Sume's STT burn covers speech only, so here is how to author the full cue list and burn it.
- caption_no_speech on a silent clip: burn Halloween text with cues
A silent AI clip fails video captions with caption_no_speech. Pass cues with text, start and end instead: a worked giveaway announcement on Sume at $0.20.
- Captions unreadable on busy footage: dim the clip, then burn
Busy B-roll can swallow burned captions. Sume can dim the clip with video-filter, then burn captions with colour overrides, for about $0.22 a clip.
Written by Sume