Flat Sume output schema for a spreadsheet row: null is an empty cell

Design a Sume output_schema that maps to one spreadsheet row. Optional fields become nullable strings, and a null becomes an empty cell when you write the CSV.

3 min readSume
All posts

How should I shape a schema that lands in a spreadsheet?

Keep it flat: one object whose properties are plain strings and numbers, one property per column. Deep nesting forces you to flatten later, and arrays of objects become several rows or a messy cell. The depth limit is 10 and the schema may have 5,000 properties, but a sheet is easier at depth one.

Sume's strict subset has one rule that surprises people. Every property must be in required. There is no truly optional field. When a value may be missing, say so with a nullable union.

How do I mark a column as optional?

Use a type union such as ["string", "null"]. The keyword nullable is rejected, and so are oneOf and allOf; anyOf is allowed. The agent returns null when it has nothing to say, and your sheet writer turns null into an empty cell.

{
  "type": "object",
  "additionalProperties": false,
  "required": ["sku", "headline", "video_url", "promo_code"],
  "properties": {
    "sku": {"type": "string"},
    "headline": {"type": "string", "maxLength": 80},
    "video_url": {"type": "string"},
    "promo_code": {"type": ["string", "null"]}
  }
}

How do I write the row?

Keep the sheet writer dumb: it should not repair data, only place it.

After the run, read output from the receipt. If it is null, the schema was not satisfied, so write the error instead. Otherwise map the keys to columns in a fixed order:

import csv, sys

COLS = ["sku", "headline", "video_url", "promo_code"]

def write_row(writer, output):
    writer.writerow(["" if output[c] is None else output[c]
                     for c in COLS])

w = csv.writer(sys.stdout)
w.writerow(COLS)
write_row(w, {"sku": "M-1", "headline": "Warm mug",
              "video_url": "https://example.com/a.mp4",
              "promo_code": None})

What else should I watch?

Keep column names stable across weeks, because renaming a property in the schema renames the key in every future output, and any formula pointing at the old column breaks silently. Add a new property and retire the old one over a release, rather than editing in place. A projection step fills some fields without seeing input, so a column that echoes a request value, such as the SKU, is safer to fill from your own row than from the agent. Join on the queue item index rather than trusting a model-returned id.

Sources

Related posts

More in Formats

All Formats posts

Written by Sume