Save a sidecar JSON with every Sume mp4: job id, model, cost, prompt

Export finished video jobs with a Python script that downloads each mp4 and writes a sidecar JSON with job id, model, cost and the prompt you saved.

5 min readSume
All posts

Write a manifest.jsonl with one row per job, then run a short Python script that polls each job, downloads the mp4 and saves a sidecar JSON next to it with the model, status, cost and your original prompt. The prompt must come from your manifest, because the poll response does not echo it back.

Why a sidecar

Sora's download URLs lasted at most an hour, according to OpenAI's guide read on 2026-10-07, and the API has been shut down since September 24, 2026. A file without its metadata is hard to reuse. A sidecar means a year from now you still know what prompt, model and spend produced a clip.

Sidecars also help when a customer asks why a clip looks the way it does. The model, the prompt and the cost sit beside the file, so support can answer from the folder without searching logs.

Field sources

Where each sidecar field comes from (read 2026-10-07)
FieldSourceNote
id, prompt, old_sora_idYour manifest.jsonlThe poll response does not return the prompt
model, statusGET /v1/videos/{id}Status must be completed
cost_usdusage.cost on the poll responseAbsent until billing is known, so the script allows null
fileLocal nameNamed after the job id

The script

Save this as sidecar.py and run it with SUME_API_KEY set. It skips jobs that are not complete, so you can rerun it until the manifest is exhausted. It uses urllib for the poll and curl for the download, and curl sends the key again because a redirect to another host drops the Authorization header.

import json, os, subprocess, urllib.request
from pathlib import Path

BASE = os.environ.get("SUME_API_BASE", "https://api.sume.com")
KEY = os.environ["SUME_API_KEY"]
OUT = Path("exports"); OUT.mkdir(exist_ok=True)

def get(url):
    req = urllib.request.Request(url, headers={"Authorization": f"Bearer {KEY}"})
    with urllib.request.urlopen(req, timeout=30) as r:
        return json.load(r)

for line in Path("manifest.jsonl").read_text().splitlines():
    row = json.loads(line)  # {"id": "job_...", "prompt": "...", "old_sora_id": "..."}
    job = get(f"{BASE}/v1/videos/{row['id']}")
    if job["status"] != "completed":
        print(row["id"], job["status"]); continue
    mp4 = OUT / f"{row['id']}.mp4"
    subprocess.run(["curl", "-sSfL", "-H", f"Authorization: Bearer {KEY}",
                    "-o", str(mp4), job["unsigned_urls"][0]], check=True)
    sidecar = {**row, "model": job["model"], "status": job["status"],
               "cost_usd": job.get("usage", {}).get("cost"), "file": mp4.name}
    mp4.with_suffix(".json").write_text(json.dumps(sidecar, indent=2))

The manifest

A manifest row looks like one line of JSON with your job id, the prompt and, if you have it, the Sora id you are replacing. The download follows unsigned_urls[0], the same link as GET /v1/videos/{id}/content?index=0, which redirects to the stored artifact.

Run the script on a schedule while you migrate. Jobs that were still rendering on the first run are picked up on the next, and the script prints the status of anything it skips so you can see what is outstanding.

After it runs

Both routes are listed in the video generation docs. Put exports/ somewhere your own backup already covers, and keep a copy of manifest.jsonl with it. The curl download post covers the one-off version of this step.

Add a checksum field if your archive is audited. Computing a SHA-256 of the mp4 after download and writing it into the sidecar takes three lines and lets you detect a corrupted copy years later.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume