chat-latest moved on Oct 7: pinning the model in a nightly job

OpenAI refreshed its chat-latest snapshot on Oct 7. Sume Agent Completions take only sume-agent. What that means for a CI job that must give stable output.

5 min readSume
All posts

OpenAI's changelog for October 7, 2026 says the chat-latest snapshot was refreshed to reflect the newest ChatGPT model, and that the docs recommend the GPT-6 model family for production API use. A floating alias like that is useful for trying the newest behavior, and risky in a CI job that compares today's output with yesterday's.

Sume's Agent Completions endpoint avoids one version of this problem by having one value: the model field accepts only sume-agent, and omitting it gives the same agent. Anything else returns 400 invalid_request. That removes model drift from your request, but it does not freeze the agent itself, and the docs do not promise a pinned build.

What you control in an Agent Completion

Because the model is fixed, the knobs that are left are the prompt, the data and the contract. Put the task in instruction or messages, put caller data in input (Sume writes all of it to /workspace/inputs/sume-action-input.json and treats it as data, not instructions), and put the result shape in output_schema.

The schema is the strongest stabilizer. The run's output is parsed against it after completion, so a CI check can fail on a missing field instead of on a wording change. If the schema cannot be satisfied, the receipt carries the reason and the run is not treated as clean output.

Request fields that reduce drift, as of 2026-10-09 (docs.sume.com/agents/completions)
FieldEffect
modelOnly sume-agent; anything else is 400 invalid_request
output_schemaBinds output to your schema
primary_output_keyNames the headline key in output
inputCaller data, written to a workspace file, not read as instructions
generation_spend_cap_usdRequired; caps generation spend for the run

A CI pattern that tolerates drift

Treat the run as a function with a contract, not a transcript. Store the receipt, assert on the schema, and diff only the fields you care about. When wording varies but the fields hold, the job passes.

  • Use a new Idempotency-Key per CI run id, so a re-run of the same pipeline returns the original receipt instead of billing a second run.
  • Set generation_spend_cap_usd low. A failing run should cost a small, known amount.
  • Archive id, thread_id and usage from the receipt with the build artifacts.
  • Pin your own prompt text in the repo. The prompt is the only part of the request that you version.

What not to assume

Do not assume that two runs a week apart used the same agent build. The docs describe the runtime (sandbox, tools, MCP bridge, media generation) but not a version. If you need reproducible output, assert on structure and keep a human review step for anything published. The aliasing lesson from OpenAI applies either way: know which parts of the request are fixed and which float.

Checking the contract in a test

A cheap test: run the same instruction with a small cap on each build, parse the receipt, and assert three things. The status is completed, the output matches the schema, and usage is not null and sits under your budget. Keep a short golden file of expected fields, not expected text.

If the test fails only on wording, tighten the schema or the prompt, not the assertion. If it fails on structure, treat it as a regression in your prompt or in the agent, open a support request with the run id, and keep the previous known-good receipt for comparison.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume