Save a Sume artifact atomically: write a .part file, then rename

A half-written MP4 that looks finished is worse than none. Download a Sume artifact to a .part file, check the length, rename once. Tested in Python.

5 min readSume
All posts

Download a finished Sume artifact to <id>.mp4.part, confirm the byte count, fsync, and only then rename it to <id>.mp4. The rename is atomic on the same filesystem, so any process that sees the final name sees a complete file. A crash, a restart or a dropped connection leaves only a .part that the next run deletes and retries. This is the cheap way to avoid shipping a truncated video that plays for six seconds and stops.

The artifact list comes from the job result and from a job.completed webhook: Sume webhooks lists payload.artifacts entries with an id, url, type and content_type. Key your files by the artifact id, not by a timestamp, and the whole thing becomes idempotent.

Why not write to the final name?

Retries are the reason this matters. A webhook can be delivered more than once, and your handler may run twice at the same moment. If both runs write straight to the final path, the second can overwrite a good file with a partial one. With the .part pattern, a second run sees the final file already exists and returns. Two overlapping runs would both write the same .part and corrupt it, so name it with the process id or use an exclusive create to avoid the collision.

What does the download code look like?

import http.server, os, tempfile, threading, urllib.request
DATA = b"x" * 300_000
class H(http.server.BaseHTTPRequestHandler):
    def do_GET(self):
        self.send_response(200)
        self.send_header("Content-Length", str(len(DATA)))
        self.end_headers()
        self.wfile.write(DATA)
    log_message = lambda *a: None
def save(url, dest):
    if os.path.exists(dest):
        return "exists"
    part = dest + "." + str(os.getpid()) + ".part"
    with urllib.request.urlopen(url, timeout=30) as r, open(part, "wb") as f:
        want = int(r.headers["Content-Length"])
        got = f.write(r.read())
        if got != want:
            os.remove(part)
            raise IOError("short read %d of %d" % (got, want))
        f.flush()
        os.fsync(f.fileno())
    os.replace(part, dest)
    return "saved"
srv = http.server.HTTPServer(("127.0.0.1", 0), H)
threading.Thread(target=srv.serve_forever, daemon=True).start()
url = "http://127.0.0.1:%d/a.mp4" % srv.server_port
dest = os.path.join(tempfile.mkdtemp(), "art_1.mp4")
assert save(url, dest) == "saved" and save(url, dest) == "exists"
assert open(dest, "rb").read() == DATA
print("atomic save ok")

How do I know the file is whole?

The example compares the received length with Content-Length, which catches a connection that drops mid-body. For a stronger check, compare against the size or checksum your own record holds, if the result gives you one. Sume jobs and results covers fetching the result; read the artifact fields there, and Sume documents media.sume.com URLs on Format run receipts as durable, but anyone holding a URL can open it, so copy the file into your own storage if you need access control, and keep the artifact id, content type and byte size in your own record so a later audit can tell a complete file from a damaged one without downloading it again.

Failure modes and what the .part pattern does, from Sume webhooks and Python os.replace (read 2026-10-06)
EventWithout .partWith .part
Crash mid-downloadTruncated file at final pathOrphan .part, retried
Duplicate webhookSecond write may clobberSecond run sees the file and exits
Short readTruncated file keptLength check deletes the part
Reader polls the folderSees a half fileSees only complete files

What about leftovers and object storage?

Clean up as part of boot: delete any *.part older than an hour before the worker starts. Do not delete fresh ones, because another process may still be writing them. On object storage the same idea is a multipart upload that you complete only after the length matches. Whichever store you use, the order stays the same: write elsewhere, verify, then publish under the final name. Acknowledge the webhook first and do the copy in the background, so your handler stays inside the delivery timeout. The example reads the body in one call to stay short; for a large clip, stream it in chunks and count bytes as you write, which keeps memory flat. Also record the final path next to the job id in your database after the rename succeeds, not before, so a row never points at a file that is not there yet.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume