Amazon A+ video descriptions per locale vs Sume caption language
A+ updateMedia upserts video descriptions by locale, one of title or descriptions per call. Sume's caption language is only a speech-to-text hint.

Amazon's updateMedia writes accessibility descriptions for a video one locale at a time: only the locales you send change, and each request updates either title or descriptions, never both. Sume's caption job is separate: it burns text onto the clip, and its language field only hints speech-to-text.
Amazon behavior is from its create, edit and retrieve A+ media page, read 2026-10-01; caption behavior from Sume's video captions docs.
What does updateMedia change?
The release notes list updateMedia among the Sep 30, 2026 additions, for updating media titles and accessibility descriptions used by Premium video modules. The page splits the cases by whether you pass associatedMediaId.
| Goal | associatedMediaId | Send |
|---|---|---|
| Video-image pairing title | As query parameter | title |
| Video-level descriptions | Omit | descriptions |
| Standalone image title | Omit | title |
How do locales behave?
Descriptions are upserted by locale. Locales that are not in the request keep what they had. To add a German description to a video that has an English one, send only the German entry. To change both, send both, in one descriptions request.
What does the Sume caption language do?
POST /v1/video-captions takes a public HTTPS video_url and burns captions on it. The language field is a speech-to-text hint (ko, en and so on); omit it for automatic detection. It does not pick the caption style or font, and it does not write an Amazon description. A silent clip fails as caption_no_speech, so for clips without speech pass authored cues with text, start and end and Sume burns that copy without speech-to-text.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: aplus-caption-de-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/example/clip.mp4",
"language": "de",
"style": "punch"
}'Are burned captions the same as accessibility descriptions?
No, they are different things. Burned captions are pixels in the file; an A+ description is metadata you send to Amazon in a given locale. Plan both: render one captioned file per language you want on screen, and send each language's description text with updateMedia. For how the language hint behaves, see language is a speech-to-text hint.
Sources
Related posts
More in Developers
- Nova Reel start_async_invoke S3 output vs a Sume result URL
Nova Reel's start_async_invoke writes output.mp4 to your S3 bucket and needs IAM. A Sume job returns a result_url and durable media.sume.com links.
- Anam 99.9% uptime SLA vs Sume failed-job retry metadata and refunds
Anam states a 99.9% uptime SLA for Cara-4. Sume documents what a failed avatar job exposes, category and retryability, and refunds the reservation.
- ant apply agent file: declare the Sume server, keep the key out
In an ant apply agent file, declare the Sume server by name and URL, https://mcp.sume.com/mcp. The API key belongs in a vault, never in the committed file.
- Argil audio upload 50 MB vs Sume Fabric's 10 MB audio URL
Argil takes mp3, wav and m4a uploads up to 50 MB. Sume's Fabric route takes a Sume-hosted audio URL up to 10 MB; H3 Max lip sync wants 5-14.8 seconds.
Written by Sume