chat-latest moved on Oct 7: pinning the model in a nightly job
OpenAI refreshed its chat-latest snapshot on Oct 7. Sume Agent Completions take only sume-agent. What that means for a CI job that must give stable output.

OpenAI's changelog for October 7, 2026 says the chat-latest snapshot was refreshed to reflect the newest ChatGPT model, and that the docs recommend the GPT-6 model family for production API use. A floating alias like that is useful for trying the newest behavior, and risky in a CI job that compares today's output with yesterday's.
Sume's Agent Completions endpoint avoids one version of this problem by having one value: the model field accepts only sume-agent, and omitting it gives the same agent. Anything else returns 400 invalid_request. That removes model drift from your request, but it does not freeze the agent itself, and the docs do not promise a pinned build.
What you control in an Agent Completion
Because the model is fixed, the knobs that are left are the prompt, the data and the contract. Put the task in instruction or messages, put caller data in input (Sume writes all of it to /workspace/inputs/sume-action-input.json and treats it as data, not instructions), and put the result shape in output_schema.
The schema is the strongest stabilizer. The run's output is parsed against it after completion, so a CI check can fail on a missing field instead of on a wording change. If the schema cannot be satisfied, the receipt carries the reason and the run is not treated as clean output.
| Field | Effect |
|---|---|
| model | Only sume-agent; anything else is 400 invalid_request |
| output_schema | Binds output to your schema |
| primary_output_key | Names the headline key in output |
| input | Caller data, written to a workspace file, not read as instructions |
| generation_spend_cap_usd | Required; caps generation spend for the run |
A CI pattern that tolerates drift
Treat the run as a function with a contract, not a transcript. Store the receipt, assert on the schema, and diff only the fields you care about. When wording varies but the fields hold, the job passes.
- Use a new
Idempotency-Keyper CI run id, so a re-run of the same pipeline returns the original receipt instead of billing a second run. - Set
generation_spend_cap_usdlow. A failing run should cost a small, known amount. - Archive
id,thread_idandusagefrom the receipt with the build artifacts. - Pin your own prompt text in the repo. The prompt is the only part of the request that you version.
What not to assume
Do not assume that two runs a week apart used the same agent build. The docs describe the runtime (sandbox, tools, MCP bridge, media generation) but not a version. If you need reproducible output, assert on structure and keep a human review step for anything published. The aliasing lesson from OpenAI applies either way: know which parts of the request are fixed and which float.
Checking the contract in a test
A cheap test: run the same instruction with a small cap on each build, parse the receipt, and assert three things. The status is completed, the output matches the schema, and usage is not null and sits under your budget. Keep a short golden file of expected fields, not expected text.
If the test fails only on wording, tighten the schema or the prompt, not the assertion. If it fails on structure, treat it as a regression in your prompt or in the agent, open a support request with the run id, and keep the previous known-good receipt for comparison.
Sources
Related posts
More in Developers
- Cheapest legal Wan 3.0 call: 2 s at 480p is $0.125, 15x less than 30 s
Wan 3.0 runs 2 to 30 s. The minimum call costs $0.125 at 480p, $0.25 at 720p and $0.50 at 1080p. A curl body, and why client-side length checks beat a 400.
- Check a GPT Image 2.5 size before you call it: a Python validator
Sume rejects GPT Image 2.5 sizes that break four rules: multiples of 16, edge 3840, ratio 3:1, 655,360 to 8,294,400 pixels. A Python check, run on 8 sizes.
- Check your image cost table against Sume /v1/images/models cost_usd
Read pricing lines from GET /v1/images/models/{id}/endpoints and compare cost_usd with your own table before a batch. Sume lines already include its margin.
- CI guard: fail the build when a new Sume pricing_skus key appears
A node:test file fetches GET /v1/videos/models and fails when a model uses a pricing_skus key your cost code cannot total. Wan, Omni and H3 rates in the table.
Written by Sume