SHA-256 manifest for a batch of 30-second clips: spot bad downloads

After downloading many 30-second Sume clips, write a manifest of size and SHA-256 per job id so a rerun skips good files and re-fetches bad ones. Python.

5 min readSume
All posts

Write a manifest line, with the job id, byte size, and SHA-256, for every clip you download from /v1/videos/{id}/content?index=0, and treat any file missing from the manifest as unfinished. A rerun then skips clips that are already good and fetches only the rest, so a crash at clip 37 of 50 does not cost you another $17.34 render.

The manifest guards your copy, not Sume's render. It cannot say a clip is the right clip, only that the bytes you saved are the bytes you finished writing.

Why a manifest and not just file existence

A 30-second MP4 is a big file, and downloads can end early when a connection drops. A file that exists with the right name may be cut short. If your rerun logic is if os.path.exists(path): skip, a truncated file survives forever and surfaces later as a player error.

The fix is to write the file under a temporary name, hash it as you go, rename it only when the stream ends cleanly, and then append to the manifest. Interrupted downloads leave a .part file and no manifest line, so they are retried.

What each state means in the manifest scheme, Sume content endpoint read 2026-10-05
State on diskManifest lineRerun does
clip.mp4 and a lineyesSkip
clip.mp4.part onlynoDownload again
clip.mp4 but no linenoHash it, add the line, or re-download
NothingnoDownload

The code

This script reads job ids from ids.txt, one per line, uses SUME_API_KEY, and downloads with the standard Authorization header. It runs the blocking part in threads so that several clips copy at once.

import asyncio, hashlib, json, os, urllib.request
from pathlib import Path

KEY = os.environ["SUME_API_KEY"]
OUT = Path("out"); OUT.mkdir(exist_ok=True)
MAN = OUT / "manifest.jsonl"
done = {json.loads(l)["job"] for l in MAN.read_text().splitlines()} if MAN.exists() else set()

def fetch(job):
    req = urllib.request.Request(
        f"https://api.sume.com/v1/videos/{job}/content?index=0",
        headers={"Authorization": "Bearer " + KEY})
    part, h, n = OUT / f"{job}.mp4.part", hashlib.sha256(), 0
    with urllib.request.urlopen(req, timeout=60) as r, open(part, "wb") as f:
        while chunk := r.read(1 << 20):
            f.write(chunk); h.update(chunk); n += len(chunk)
    part.rename(OUT / f"{job}.mp4")
    return {"job": job, "bytes": n, "sha256": h.hexdigest()}

async def main():
    jobs = [j for j in Path("ids.txt").read_text().split() if j not in done]
    for coro in asyncio.as_completed([asyncio.to_thread(fetch, j) for j in jobs]):
        try:
            line = await coro
        except Exception as e:
            print("failed:", e); continue
        with MAN.open("a") as m: m.write(json.dumps(line) + "\n")

asyncio.run(main())

Using it with a batch

Run the script after each wave of completed jobs, not just once at the end. Compare the manifest against your list of expected ids; any id missing from both the manifest and your failure log needs attention. If you also store clips in object storage, upload from the manifest, and keep the hash as object metadata so a later audit can compare.

Remember the second bill: a failed download is free to retry, but a failed render is a different matter. Check the job status first, and only re-fetch completed jobs.

What the hash does not prove

A SHA-256 recorded at download time proves that a later copy of the file is identical to the one you saved. It does not prove the first download was complete, because a stream that ends early hashes just as cleanly as a full one. For that you rely on the clean end of the response and on a size sanity check.

A simple extra guard is to flag any clip whose size is far below the others in the same batch. All the clips in a batch that share a model, resolution and length should fall in a similar range, so an outlier is worth a second look before it goes to an editor.

Keep the manifest next to the clips, commit it to your own store, and treat it as the source of truth for what is finished.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume