What to store from a Sume result: job and run artifact columns

Store artifact id, durable media.sume.com URL, type and content type for jobs, and size, width, height, duration and sha256 for runs. SQLite schema included.

5 min readSume
All posts

Store the artifact id, the URL, the type and the content type from every result, and add size_bytes, width, height, duration_ms and checksum_sha256 when the result comes from a run. A job artifact is {id, url, type, content_type}. A run artifact has those four fields plus the five extra ones. The URL is a durable public link on media.sume.com, so you can keep it instead of copying the file.

Decide this before the first production run. Teams that store only the URL find, months later, that they cannot answer basic questions such as how large a file was, how long a clip ran or whether a copy still matches the original.

Why the run shape has more fields

The two shapes differ because they answer different needs. A job result is a handle on one generated file. A run receipt is the record of a finished Format, Action or Agent run, and it carries enough metadata to verify the file without downloading it first. The checksum is the part many teams skip, and it is the one that lets a nightly job prove that a stored file is the one Sume produced.

This is also why the receipt fields line up with typical delivery rules. A destination that wants a minimum size, a certain aspect ratio or a maximum duration can be checked from the stored columns, with no download and no call to a media tool.

One table for both shapes

Use one table for both and let the run only columns be null for jobs.

Artifact fields by result type (read 2026-10-05)
ColumnJob resultRun receipt
idYesYes
urlYes, durableYes, durable
typeYesYes
content_typeYesYes
size_bytesNoYes
width, heightNoYes
duration_msNoYes
checksum_sha256NoYes

The schema

The statement below creates that table and inserts both kinds. Keep the source id, either job_ or arun_, in its own column, because the two families are polled at different routes. Dollar amounts do not belong in this table. Keep them in a separate ledger keyed by the same id, using the usage.billable_amount_usd_micros value from the receipt.

Using one table keeps queries simple. Joining a job artifact with a run artifact by artifact_id or by source_id needs no special case, and a report on stored media by type or by size works across both families.

import sqlite3

db = sqlite3.connect(":memory:")
db.execute('''create table artifacts (
  artifact_id text primary key, source_id text not null,
  url text not null, type text, content_type text,
  size_bytes integer, width integer, height integer,
  duration_ms integer, checksum_sha256 text)''')

job_art = {"artifact_id": "artf_1", "url": "https://media.sume.com/artifacts/a", "type": "image", "content_type": "image/png"}
run_art = dict(job_art, artifact_id="artf_2", size_bytes=1024, width=1080, height=1920, duration_ms=6000, checksum_sha256="ab" * 32)

for src, a in (("job_demo", job_art), ("arun_demo", run_art)):
    cols = ", ".join(a); marks = ", ".join("?" * (len(a) + 1))
    db.execute(f"insert into artifacts (source_id, {cols}) values ({marks})", [src, *a.values()])
print(db.execute("select source_id, type, width from artifacts").fetchall())

Which URLs are safe to keep

Store the URL, and read the result again only when you need fresh metadata. Media URLs on media.sume.com are durable and public, which is different from the download URL of the asset library, a short lived presign. The asset routes are also hidden from the public OpenAPI document, so do not store those links as if they were permanent.

If you do need a longer lived copy under your own control, download the file once, check it, and store it in your bucket, while keeping the Sume URL as the reference. Never store a signed or short lived link as the only record of a result. Doing so leaves you with a dead reference after the link expires.

Verify after copying

Check the checksum after you copy a file anywhere else. Hash the downloaded bytes with SHA-256 and compare the hex digest with checksum_sha256. A mismatch means a truncated or altered copy, so download again and do not mark the order as delivered. For jobs, which carry no checksum, record your own hash on first download, so later copies can be compared with it.

Keep the checksum next to the URL in your audit log, too. It costs 64 characters and it makes a later dispute about which file was delivered easy to settle, because both sides can compare the same hash.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume