Video analyses: include_transcript false leaves scene audio null
On the legacy Sume video-analyses resource, scene audio is null unless include_transcript was true. Handle both shapes, and know what replaces it for new work.

If you still read Sume's legacy video_analysis resource, the audio field on each scene is null unless the job was created with include_transcript: true. With it on, each scene's audio is { speech, has_speech, has_music }. A parser that does scene["audio"]["speech"] crashes on every scene of an analysis that was created without it.
Retiring a media surface is a normal event now; OpenAI published a Help Center article titled "What to know about the Sora discontinuation" (OpenAI Help Center). Sume's own legacy surface has a documented status too, which is the first thing to check before you write new code against it.
Where this surface stands
Per the video analyses docs, on dest POST /v1/video-analyses answers 410 video_analysis_retired, while production accepts create until the follow-up work tracked as #5953 PR-C2, and stored vana_ rows stay readable by GET in both environments. New clip inspection is video inspect: probe facts, stills and an optional transcript, with no typed scenes. Semantic questions on dest are video_analyze or video_segment, only where those names appear in tools_list.
| Created with | scene.audio | Where speech text is |
|---|---|---|
| include_transcript: true | { speech, has_speech, has_music } | scene.audio.speech |
| include_transcript omitted or false | null | Not requested; no per-scene speech |
Read both shapes
Reading stored rows is the case that still matters on every environment. Treat audio as optional and branch once.
scenes = [
{"start_seconds": 0, "end_seconds": 4,
"audio": {"speech": "Welcome back", "has_speech": True, "has_music": False}},
{"start_seconds": 4, "end_seconds": 9, "audio": None},
]
for scene in scenes:
audio = scene.get("audio")
if audio is None:
line = "no transcript requested"
elif audio["has_speech"]:
line = audio["speech"]
else:
line = "music only" if audio["has_music"] else "silent"
span = f"{scene['start_seconds']}-{scene['end_seconds']}s"
print(span, "|", line)
Do not start new work here. The docs also note the accuracy caveat that longer videos are less accurate and short clips are the most reliable, and that each accepted analysis reserved $0.30 under the current fixed estimate. See the 410 post for the migration.
Sources
Related posts
More in Media tools
- Video analyses keyframes: one still per second per scene
A legacy Sume video analysis gives a still for each whole second of a scene plus a keyframe_url near 40%. How stills fail and how to pick your own cover.
- Video filter /check: submit or fix and recheck
POST /v1/video-filter/check is unbilled and returns valid, diagnostics and a next_action of submit_video_filter or fix_program_and_recheck. Branch on it.
- Video inspect fast seek: requested_times vs sample_times
A fast-seek inspect grid returns requested_times and sample_times. Use sample_times for what the tiles show, then trim from them, never from the request.
- WCAG 1.4.4 resize text: do burned-in captions need it?
WCAG 1.4.4 exempts captions and images of text, so burned-in captions need no resize. You still choose their size: use design.typography on Sume.
Written by Sume