SHA-256 manifest for a batch of 30-second clips: spot bad downloads
After downloading many 30-second Sume clips, write a manifest of size and SHA-256 per job id so a rerun skips good files and re-fetches bad ones. Python.

Write a manifest line, with the job id, byte size, and SHA-256, for every clip you download from /v1/videos/{id}/content?index=0, and treat any file missing from the manifest as unfinished. A rerun then skips clips that are already good and fetches only the rest, so a crash at clip 37 of 50 does not cost you another $17.34 render.
The manifest guards your copy, not Sume's render. It cannot say a clip is the right clip, only that the bytes you saved are the bytes you finished writing.
Why a manifest and not just file existence
A 30-second MP4 is a big file, and downloads can end early when a connection drops. A file that exists with the right name may be cut short. If your rerun logic is if os.path.exists(path): skip, a truncated file survives forever and surfaces later as a player error.
The fix is to write the file under a temporary name, hash it as you go, rename it only when the stream ends cleanly, and then append to the manifest. Interrupted downloads leave a .part file and no manifest line, so they are retried.
| State on disk | Manifest line | Rerun does |
|---|---|---|
| clip.mp4 and a line | yes | Skip |
| clip.mp4.part only | no | Download again |
| clip.mp4 but no line | no | Hash it, add the line, or re-download |
| Nothing | no | Download |
The code
This script reads job ids from ids.txt, one per line, uses SUME_API_KEY, and downloads with the standard Authorization header. It runs the blocking part in threads so that several clips copy at once.
import asyncio, hashlib, json, os, urllib.request
from pathlib import Path
KEY = os.environ["SUME_API_KEY"]
OUT = Path("out"); OUT.mkdir(exist_ok=True)
MAN = OUT / "manifest.jsonl"
done = {json.loads(l)["job"] for l in MAN.read_text().splitlines()} if MAN.exists() else set()
def fetch(job):
req = urllib.request.Request(
f"https://api.sume.com/v1/videos/{job}/content?index=0",
headers={"Authorization": "Bearer " + KEY})
part, h, n = OUT / f"{job}.mp4.part", hashlib.sha256(), 0
with urllib.request.urlopen(req, timeout=60) as r, open(part, "wb") as f:
while chunk := r.read(1 << 20):
f.write(chunk); h.update(chunk); n += len(chunk)
part.rename(OUT / f"{job}.mp4")
return {"job": job, "bytes": n, "sha256": h.hexdigest()}
async def main():
jobs = [j for j in Path("ids.txt").read_text().split() if j not in done]
for coro in asyncio.as_completed([asyncio.to_thread(fetch, j) for j in jobs]):
try:
line = await coro
except Exception as e:
print("failed:", e); continue
with MAN.open("a") as m: m.write(json.dumps(line) + "\n")
asyncio.run(main())Using it with a batch
Run the script after each wave of completed jobs, not just once at the end. Compare the manifest against your list of expected ids; any id missing from both the manifest and your failure log needs attention. If you also store clips in object storage, upload from the manifest, and keep the hash as object metadata so a later audit can compare.
Remember the second bill: a failed download is free to retry, but a failed render is a different matter. Check the job status first, and only re-fetch completed jobs.
What the hash does not prove
A SHA-256 recorded at download time proves that a later copy of the file is identical to the one you saved. It does not prove the first download was complete, because a stream that ends early hashes just as cleanly as a full one. For that you rely on the clean end of the response and on a size sanity check.
A simple extra guard is to flag any clip whose size is far below the others in the same batch. All the clips in a batch that share a model, resolution and length should fall in a similar range, so an outlier is worth a second look before it goes to an editor.
Keep the manifest next to the clips, commit it to your own store, and treat it as the source of truth for what is finished.
Sources
Related posts
More in Developers
- Shorts series episodes number by publish date: a Python order check
YouTube numbers Shorts series episodes by publish date, so upload order is episode order. Check durations and publish times in Python before you schedule.
- Shorts series: one Idempotency-Key per season and episode
Timeline renders require an Idempotency-Key. Name it from season and episode, so a retried upload script for episode 4 does not queue a second render.
- Shrink an AI image to a byte budget: JPEG quality search in Pillow
Platforms cap image files at 1 MB or 5 MB. Binary-search the JPEG quality in Pillow for the sharpest file under your byte limit, from a Sume image.
- silence_split_seconds: tune caption line breaks from STT segments
Sume STT sentence segmentation can split on silence. silence_split_seconds takes 0.2 to 3 and returns gapless segments you can use as caption lines.
Written by Sume