GPT-6.1 Sol function tool for Sume runs: clamp the cap in code

A Responses API function tool that lets GPT-6.1 Sol start a Sume Agent Completion: call_id as Idempotency-Key and a spend cap the model cannot exceed.

6 min readSume
All posts

To let GPT-6.1 Sol start a Sume Agent Completion, define a Responses API function tool, run its function_call in your own code, and send back a function_call_output. In the handler, clamp generation_spend_cap_usd to a ceiling you set, use the call_id as the Idempotency-Key, and return the run id, not a finished video. The model chooses the task; your code owns the money and the key.

The OpenAI side

OpenAI's function calling guide describes a function tool with type, name, description, parameters and strict. The model returns function_call items with a call_id, a name and JSON arguments; you reply with a function_call_output carrying the same call_id and an output. GPT-6.1 Sol is priced at $2 per million input tokens and $10 per million output tokens, with cached input at $0.10, per the OpenAI community post.

The Sume side

Agent Completions accept instruction or messages, require generation_spend_cap_usd, and answer 202 with an agent.run receipt containing id, status and status_url. Replaying an Idempotency-Key returns the original receipt, which is exactly what you want when a model retries a tool call.

import json, os, urllib.request

MAX_CAP = 5.0

def start_run(args: dict, call_id: str) -> dict:
    cap = min(float(args["generation_spend_cap_usd"]), MAX_CAP)
    if not 0 < cap <= MAX_CAP:
        raise ValueError("cap must be positive")
    body = json.dumps({"instruction": args["instruction"],
                       "generation_spend_cap_usd": cap}).encode()
    req = urllib.request.Request(
        "https://api.sume.com/v1/agent/completions", data=body, method="POST",
        headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
                 "Content-Type": "application/json",
                 "Idempotency-Key": call_id})
    with urllib.request.urlopen(req, timeout=30) as resp:
        run = json.load(resp)["data"]
    out = {"run_id": run["id"], "status": run["status"],
           "status_url": run["status_url"]}
    return {"type": "function_call_output", "call_id": call_id,
            "output": json.dumps(out)}

Why the clamp lives in code

Strict mode makes the arguments well-typed, not wise. If the schema allows any number, a model can ask for a 50 dollar cap; min() in the handler makes that request harmless. Sume's docs describe the cap as the substitute for the interactive spend-approval prompt that a backend caller does not have, so treat it as policy you own.

Who decides what (read 2026-10-04)
DecisionOwnerMechanism
What to makeGPT-6.1 Solinstruction argument
How much may be spentYour handlermin() against MAX_CAP
Whether a retry pays twiceSumeIdempotency-Key replay
Which key is usedYour serverSUME_API_KEY env var

Then poll, do not wait

The tool result should return in seconds, even though the run may take minutes. Have a second tool, or your own background worker, read status_url until next_action stops being poll_status, or use a run webhook. Keep the key out of logs and tool outputs, per Safe automation.

One caveat: the key needs the agent_completions:write scope, and a key created before Agent Completions shipped does not have it. Create a new one in the dashboard, as described on Authentication.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume