GPT-6 Sol tool calls: use the Responses API to drive Sume

GPT-6 Sol only calls functions in Chat Completions when reasoning effort is none. For Sume's MCP tools, use the Responses API and keep reasoning on.

4 min readSume
All posts

Drive Sume from GPT-6 Sol through the Responses API. OpenAI's model page says function calling works in Chat Completions only when reasoning effort is none, and that the Responses API gives the best performance. A Sume workflow plans, waits and retries, which is the work reasoning is for.

The Sol facts are from OpenAI's model page, read 2026-10-01. The Sume facts are from its MCP docs.

What does the Sol page say?

It lists a 1,050,000-token context, $2 input and $10 output per 1M tokens, and the Chat Completions, Responses and Batch endpoints. Reasoning effort accepts none, low, medium (the default), high, xhigh and max.

GPT-6 Sol tool use, from OpenAI's page, read 2026-10-01.
EndpointFunction callingReasoning
Chat CompletionsOnly when effort is noneOff for tool turns
ResponsesSupportedAny effort; recommended

Why does that matter for Sume?

A single Sume job is several steps: preview with dry_run, submit with an idempotency_key, call jobs_wait in slices and, on wait_slice_expired, wait again with the same ids. Per the tools and gates docs, jobs_wait defaults to 50 seconds and caps at 55. A model with reasoning off can follow that loop, but it has no room to decide when a result needs a second look.

Can Chat Completions still work?

Yes, with effort set to none and your own function definitions that call Sume's REST or MCP endpoints. Keep the tools list small, and put the retry rule in the tool description: a 524 is transport, never a job outcome. If you need reasoning between calls, move to Responses.

Sources

Related posts

More in Integrations

All Integrations posts

Written by Sume