Video analyses: include_transcript false leaves scene audio null

On the legacy Sume video-analyses resource, scene audio is null unless include_transcript was true. Handle both shapes, and know what replaces it for new work.

5 min readSume
All posts

If you still read Sume's legacy video_analysis resource, the audio field on each scene is null unless the job was created with include_transcript: true. With it on, each scene's audio is { speech, has_speech, has_music }. A parser that does scene["audio"]["speech"] crashes on every scene of an analysis that was created without it.

Retiring a media surface is a normal event now; OpenAI published a Help Center article titled "What to know about the Sora discontinuation" (OpenAI Help Center). Sume's own legacy surface has a documented status too, which is the first thing to check before you write new code against it.

Where this surface stands

Per the video analyses docs, on dest POST /v1/video-analyses answers 410 video_analysis_retired, while production accepts create until the follow-up work tracked as #5953 PR-C2, and stored vana_ rows stay readable by GET in both environments. New clip inspection is video inspect: probe facts, stills and an optional transcript, with no typed scenes. Semantic questions on dest are video_analyze or video_segment, only where those names appear in tools_list.

Which shape you get (video analyses docs, read 2026-10-04)
Created withscene.audioWhere speech text is
include_transcript: true{ speech, has_speech, has_music }scene.audio.speech
include_transcript omitted or falsenullNot requested; no per-scene speech

Read both shapes

Reading stored rows is the case that still matters on every environment. Treat audio as optional and branch once.

scenes = [
    {"start_seconds": 0, "end_seconds": 4,
     "audio": {"speech": "Welcome back", "has_speech": True, "has_music": False}},
    {"start_seconds": 4, "end_seconds": 9, "audio": None},
]

for scene in scenes:
    audio = scene.get("audio")
    if audio is None:
        line = "no transcript requested"
    elif audio["has_speech"]:
        line = audio["speech"]
    else:
        line = "music only" if audio["has_music"] else "silent"
    span = f"{scene['start_seconds']}-{scene['end_seconds']}s"
    print(span, "|", line)

Do not start new work here. The docs also note the accuracy caveat that longer videos are less accurate and short clips are the most reliable, and that each accepted analysis reserved $0.30 under the current fixed estimate. See the 410 post for the migration.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume