Save a Sume job's result artifacts in Python by content type

Fetch GET /v1/jobs/:id/result and save each artifact with an extension from content_type, not the URL. Standard library only, with a text-result guard.

4 min readSume
All posts

Read GET /v1/jobs/:id/result, loop over data.result.artifacts, and name each file from your own job id plus the extension that its content_type implies. Do not parse the URL. The OpenAPI spec says artifact URLs are opaque Sume media CDN paths and clients must not parse them. Guard for results with no artifacts: a speech-to-text job returns text and optional word timings, not media.

What an artifact holds

Each entry in result.artifacts has four required fields. The URL points to the Sume media CDN, and the visibility field, when present, is public_by_link, so the download needs no API key.

PublicArtifact fields from the OpenAPI spec, read 2026-10-08
FieldUse it for
idAn opaque artifact id. Safe to log. Treat it as a token.
typeThe kind of media, such as image
urlWhere to download the file (https://media.sume.com/artifacts/...)
content_typeThe MIME type, such as image/png. Use it for the file extension.

The script

It uses only the standard library. The mimetypes module maps a MIME type to an extension, and a fallback of .bin keeps an unknown type from crashing the run. The result endpoint answers 409 job_not_completed until the job completes, so call this after your status poll shows the job completed, not before.

import json, mimetypes, os, urllib.request

def ext(content_type):
    return mimetypes.guess_extension(content_type.split(";")[0].strip()) or ".bin"

def save_artifacts(job_id, folder="."):
    req = urllib.request.Request(
        f"https://api.sume.com/v1/jobs/{job_id}/result",
        headers={"x-api-key": os.environ["SUME_API_KEY"]})
    with urllib.request.urlopen(req, timeout=30) as r:
        result = json.load(r)["data"]["result"] or {}
    paths = []
    for i, art in enumerate(result.get("artifacts") or []):
        path = os.path.join(folder, f"{job_id}-{i}{ext(art['content_type'])}")
        urllib.request.urlretrieve(art["url"], path)  # opaque CDN link
        paths.append(path)
    return paths

if __name__ == "__main__":
    print(save_artifacts(os.environ["JOB_ID"]) or "no media artifacts (text result?)")

Details that matter

For failed or canceled jobs the result call has nothing to return, so read the error from the job record at GET /v1/jobs/:id instead.

  • Name files from the job id and the index, so a rerun overwrites the same file instead of piling up copies.
  • Read the result with the same key that created the job. Another member's key gets 404.
  • A mime type that the local table does not know falls back to .bin. Rename by hand or extend the table with mimetypes.add_type.
  • Download soon after completion and copy the file to your own storage if you need it long term.

Why not parse the URL

It is tempting to take the extension from the end of the URL. The spec warns against it: artifact paths are opaque and omit workspace ids, user ids, job ids and workflow details. The content_type field is the stable contract. If a job returns several artifacts, such as an Avatar Video with preview frames, the index in the file name keeps them apart, and type tells you which one is the main media.

If you only want the first video, filter on art['type'] before the download loop instead of saving everything.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume