Store the Sume artifact, not a provider link: what to persist per job

A completed Sume job returns artifacts with id, url, type and content_type. Persist those plus the job id and idempotency key, not provider links.

4 min readSume
All posts

Save the job id, your idempotency key and, for each artifact, its id, type and content_type, and re-read the job whenever you need a link. The Sume jobs doc says to use the Sume media URLs from the result and that raw provider URLs are not public API outputs, so there is nothing else worth saving.

A completed job looks like result.artifacts[], each entry holding id, url, type and content_type. Fields below come from the jobs and results page, read 2026-10-10.

A table to model

What to store per completed job, from the jobs doc read 2026-10-10
FieldStore it?Why
Job idYes, as the primary referenceLets you re-read status, result and usage
Your Idempotency-KeyYes, written before the submitJoins your row to the job after a crash
Artifact idYesStable name for the file, independent of its link
Artifact type and content_typeYesPick the player or the viewer without guessing from the URL
Artifact urlCache onlyTreat it as a link to re-fetch, not as a permanent address
Provider URL or internal model idNoNot a public output and not promised

Why not just keep the URL

The docs do not promise how long a media link works, so a design that assumes it lasts forever is making a bet the API has not signed. Keeping the job id costs a few bytes and lets you ask Sume again. If you need the file to live in your own bucket, download it right after completion and store your copy, with the artifact id as its name.

  • Download from the Sume URL, not from any URL you saw in a log or a provider tool.
  • Store content_type so that a PNG and a JPEG are never confused after a re-encode.
  • Keep the usage row id or the job id next to your invoice line so cost questions can be answered later.

Refreshing a link

On a read, call GET /v1/jobs/{id} and take the artifacts from result. If you listed jobs instead, note that list pages are newest first, so map by id and idempotency_key, not by array position. The pagination post covers the cursor loop.

A schema sketch

A single table is enough for most apps. Key it on your own order or item id, and keep the Sume fields beside it so that every question about a file can be answered from one row.

Suggested columns for a generation record, my design based on the jobs doc read 2026-10-10
ColumnHoldsWritten when
item_idYour own idBefore the submit
idempotency_keyThe header valueBefore the submit
job_idThe Sume job idAfter the POST returns
statusLatest job statusEach poll or webhook
artifact_id, type, content_typeFrom result.artifacts[]On completion
own_copy_pathWhere you saved the fileAfter your download

Mistakes to avoid

Three habits cause most of the pain. The first is parsing the file type out of the URL; the artifact already tells you its type and content_type, so read those fields. The second is saving the whole job JSON as the only record, which makes simple queries such as "all videos from last week" slow and brittle. The third is forgetting the failed and canceled jobs: store their status and error category too, so a support question about a missing file has an answer.

A short rule helps teams: if a value was returned by Sume and you need it again next month, store its id; if you need the bytes, store your own copy.

  • Index job_id and idempotency_key as unique columns.
  • Store the terminal status even when no artifact exists.
  • Download inside a background task with a retry, not inside the request that served the user.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume